Technology

Anthropic Releases Claude Opus 5: Near-Frontier Intelligence at Half the Cost of Fable 5

Martin HollowayPublished 7d ago5 min readBased on 5 sources
Reading level
Anthropic Releases Claude Opus 5: Near-Frontier Intelligence at Half the Cost of Fable 5

Anthropic announced Claude Opus 5 on July 24, 2026, positioning it as a model that approaches the frontier intelligence of Claude Fable 5 at half the price. The model is available now and serves as the new default on Claude Max and the strongest model available on Claude Pro (Anthropic).

The pricing and performance relationship to Fable 5 is the core commercial claim. On CursorBench 3.2 at max effort, Opus 5 performs within 0.5% of Fable 5's peak score at half the cost per task. On OSWorld 2.0, a computer use benchmark, Opus 5 surpasses Fable 5's best result at just over a third of the cost. The model also provides greatly improved performance over its predecessor Opus 4.8 at the same cost, more than doubling Opus 4.8's performance on Frontier-Bench v0.1 at a lower cost per task.

According to Anthropic's official documentation, the largest gains over Opus 4.8 are in deep reasoning, agentic and long-horizon tasks, and test-time compute (Anthropic Docs). Anthropic describes Opus 5 as a "thoughtful and proactive" model and has positioned it as designed for powering long-running agents (Anthropic).

On Frontier-Bench and GDPval-AA evaluations, Opus 5 is the new state-of-the-art. Several benchmark results stand out. On ARC-AGI 3, Opus 5 scores three times as high as the next-best model. On Zapier AutomationBench, its pass rate is roughly 1.5x the next-best model for the same cost per task, and at its lowest effort setting it passes more tasks than any other model. One notable exception: Opus 5 remains behind Mythos 5 on cybersecurity tasks.

The agentic and long-horizon capabilities surface in concrete task completions that earlier models could not achieve. On a Frontier-Bench task requiring reconstruction of a machine part as a 3D FreeCAD model from a drawing, Opus 5 wrote its own computer vision pipeline and succeeded repeatedly while no competing model could solve it after five attempts. In a separate case, Opus 5 found the root cause of a real bug in a popular open-source package manager and fixed an edge case that the community's own patch had missed; a competing model fixed only the surface symptom. An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session, a task previous models could not complete. Opus 5 is also the strongest Opus model Anthropic has tested on its internal trading benchmark.

In the life sciences, Opus 5 outperforms Opus 4.8 across all of Anthropic's internal evaluations covering structural biology, organic chemistry, and bioinformatics. On organic chemistry tasks involving inference of molecular structures from spectroscopy data, Opus 5 scores 10.2 percentage points higher than Opus 4.8. On protein-related tasks such as predicting how sequence variations affect function, the gap is 7.7 percentage points.

On the safety front, Anthropic's system card reports that Opus 5 is the company's most aligned model to date on its automated behavioral audit, surpassing Sonnet 5, Opus 4.8, and Mythos 5 (Anthropic System Card).

The pricing structure here is worth pausing on. A model that lands within half a percentage point of the frontier on a coding benchmark like CursorBench 3.2, while costing half as much per task, changes the economics of running autonomous coding agents in production. The OSWorld 2.0 result is arguably more striking: surpassing Fable 5's best computer-use score at roughly a third of the cost suggests that the cost-per-unit-of-agentic-work curve is bending in a way that matters for anyone building agents that interact with real desktop environments. The ARC-AGI 3 result, at 3x the next-best score, hints at gains in abstract reasoning that go beyond what benchmark-watching alone can fully contextualize.

The cybersecurity gap behind Mythos 5 is a real limitation, not a footnote. For teams evaluating models for security-sensitive agentic workflows, that gap narrows the decision space. And the life sciences gains, while framed against the prior Opus generation rather than against specialized tools, suggest the model is becoming a credible collaborator on tasks like spectroscopy interpretation and protein function prediction, where the cost of errors is high and the domain-specific vocabulary is unforgiving.

What Opus 5 enables, taken in aggregate, is a class of long-running agentic work that previous models could not complete at all, now available at a price point that makes sustained deployment plausible rather than experimental. The FreeCAD pipeline example and the single-session market data feed build are illustrative not because they are flashy, but because they represent the kind of multi-step, multi-tool problem that has been the persistent gap between benchmark scores and real utility. That gap appears to be closing.