The rapid cycle of model updates from major labs has become a defining feature of the current period in AI research and deployment.
New releases arrive with enough frequency that keeping track of incremental gains in capability, cost, and reliability requires steady attention, especially as the distinctions between tiers grow more nuanced and the practical differences for everyday users shift in subtle ways.
And this time, Anthropic released 'Claude Opus 5,' to which the company describes it as a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5, but with half the price.
It is positioned as a step change for the Opus tier, with particular strength in coding, knowledge work, and long-running agentic tasks.
On evaluations such as Frontier-Bench v0.1, Opus 5 more than doubles the performance of Opus 4.8 while operating at a lower cost per task and surpasses other models overall.
On CursorBench 3.2 it reaches within 0.5 percent of Fable 5’s peak score at maximum effort yet does so at roughly half the cost per task. It also sets new highs on several other coding and professional benchmarks while remaining behind Mythos 5 on certain cybersecurity measures.
Efficiency receives equal emphasis.
Opus 5 outperforms competing models at comparable or lower cost per task across multiple workloads. On ARC-AGI-3, an evaluation that requires solving novel problems, its score is three times higher than the next-best model.
Similar advantages appear on automation suites such as Zapier AutomationBench and on OSWorld 2.0, where it exceeds every other model at a given cost level and surpasses Fable 5’s best result at a little more than one-third the expense.
Life-sciences evaluations show consistent gains over Opus 4.8, including roughly ten percentage points higher accuracy when inferring molecular structures from spectroscopy data and nearly eight points higher when predicting protein sequence variations and function.
The model also improves on visual generation and agentic behavior.
Early examples include more coherent visualizations of airflow over aerodynamic surfaces and interactive cell diagrams.
In software engineering scenarios it has demonstrated the ability to construct complete computer-vision pipelines from limited inputs, locate root causes of deep bugs that other systems missed, and sustain multi-step workflows such as building market-data feeds with self-generated test harnesses.
Consistency metrics rise as well, with reduced variance across runs and stronger self-verification during extended tasks.
Alignment testing places Opus 5 as Anthropic’s most aligned model to date according to the company’s automated behavioral audit.
It records the lowest rates of reckless or deceptive behavior among recent Claude models and the strongest adherence to the principles set out in Claude’s Constitution.
On cybersecurity tasks it is stronger than Opus 4.8 at identifying vulnerabilities yet remains substantially behind Mythos 5 at developing exploits.
Safeguards are calibrated to permit legitimate vulnerability discovery and remediation while blocking high-risk uses such as binary scanning, penetration testing, and exploit generation; flagged requests can fall back to earlier models.
Opus 5 is available immediately on all paid Claude plans and through the Claude API at the same pricing as Opus 4.8.
It becomes the default model on Claude Max and the strongest option on Claude Pro.
A Fast mode runs at approximately 2.5 times the default speed. Effort settings remain adjustable so users can trade intelligence against token usage and latency according to the demands of a given task. No general data retention applies for ordinary access, and the model inherits the existing cyber and biology safeguards with refinements that route certain dual-use queries appropriately.
Taken together, the release narrows the practical gap between the Opus and Fable tiers for many coding, research, and enterprise workloads while preserving the cost structure of the previous Opus generation.
Independent reporting and early user feedback largely echo the benchmark patterns Anthropic published, though long-term real-world reliability and edge-case behavior will continue to be measured in the months ahead.
That said, the model is not without constraints. Anthropic notes that Opus 5 still shows important limitations on long-running autonomous research tasks, particularly in biology, where Mythos 5 remains stronger.
It also lags substantially behind Mythos 5 when it comes to developing exploits in cybersecurity, even though it has improved at identifying vulnerabilities. In the company’s system card, evaluators observed occasional unproductive self-verification loops and a tendency to produce longer, more detailed refusals that sometimes included more operational context than ideal.
Knowledge is cut off around May 2026, and the model sits in the moderate latency range compared with faster options in the same family.
Read: Anthropic's Fable 5 and Mythos 5 Are Considered a National Security Risk, Says U.S. Government
















































































































































































































































































































































































