Anthropic's New AI Model Does Almost as Well as Its Best — for Half the Price

Anthropic announced Claude Opus 5 on July 24, 2026. The company says it performs nearly as well as its most powerful model, Claude Fable 5, but costs half as much to use. It is available now as the default model on Claude Max and the strongest option on Claude Pro (Anthropic).
The key claim is about value. To test AI models, researchers use benchmarks — standardized sets of tasks that let you compare one model to another, like a standardized test for AI. On a coding benchmark called CursorBench 3.2, Opus 5 scores within 0.5% of Fable 5's top result but costs half as much per task. On OSWorld 2.0, which tests how well a model can use a computer desktop the way a person would, Opus 5 actually beats Fable 5's best score at about a third of the cost. It also more than doubles the performance of the previous version, Opus 4.8, on another benchmark called Frontier-Bench v0.1 (Anthropic Docs).
Think of benchmarks like fuel efficiency ratings for cars. They give you a rough basis for comparison, even if real-world mileage varies. Opus 5's results across several of these tests are the best any model has achieved so far.
The biggest improvements over Opus 4.8 are in tasks that require the model to work through a problem step by step over a long time, take actions on its own to reach a goal, and spend more time reasoning before answering. Anthropic calls Opus 5 a "thoughtful and proactive" model designed for long-running agents — AI programs that carry out multi-step tasks with limited human supervision (Anthropic).
On a test of abstract reasoning called ARC-AGI 3, Opus 5 scores three times as high as the next-best model. On Zapier AutomationBench, which measures how well a model can set up automated workflows across different apps, its pass rate is about 1.5 times the next-best model at the same cost. One exception: Opus 5 still trails a competitor called Mythos 5 on cybersecurity tasks.
The model's capabilities show up in real-world tasks that earlier versions could not do. In one test, Opus 5 was given a drawing of a machine part and had to rebuild it as a 3D model using FreeCAD, a free design program. It wrote its own software to analyze the image and succeeded repeatedly, while no competing model could solve the problem even after five tries. In another case, Opus 5 found the underlying cause of a real bug in a popular open-source software tool and fixed a subtle issue that the community's own repair had missed. A competing model only fixed the visible symptom.
An engineer at a trading firm used Opus 5 to build a market data feed for a new stock exchange in a single session — something previous models could not finish. Opus 5 is also the best Opus model Anthropic has tested on its internal trading benchmark.
In the life sciences, Opus 5 outperforms Opus 4.8 across all of Anthropic's internal tests covering structural biology, organic chemistry, and bioinformatics. On chemistry tasks that involve figuring out the structure of a molecule from spectroscopy data (using light-based measurements to identify what a substance is made of), Opus 5 scores 10.2 percentage points higher. On protein-related tasks, like predicting how small changes in a protein's building blocks affect what it does, the gap is 7.7 percentage points.
On the safety front, Anthropic's system card reports that Opus 5 is the company's most aligned model to date on its automated behavioral audit — essentially a test of whether the model follows safety guidelines it was given — beating out Sonnet 5, Opus 4.8, and Mythos 5 (Anthropic System Card).
The broader picture here is about cost. When a model scores nearly as well as the most expensive option on coding tasks but costs half as much, the math changes for companies that want to run AI agents continuously. The OSWorld 2.0 result — beating the top model at a third of the cost — is even more striking for anyone building systems where AI interacts with real computers. And the ARC-AGI 3 score, at three times the next-best result, suggests gains in reasoning that go beyond what the numbers alone can tell us.
The cybersecurity gap behind Mythos 5 is a real limitation, not a minor detail. For teams looking at models for security-sensitive work, that gap narrows their choices. The life sciences gains, while measured against the previous Opus version rather than against specialized scientific tools, suggest the model is becoming a useful partner on tasks like interpreting spectroscopy data and predicting protein function, where mistakes are costly and the terminology leaves little room for error.
What Opus 5 enables, taken together, is a category of long-running work that previous models could not complete at all, now available at a price that makes regular use plausible rather than experimental. The FreeCAD example and the single-session market data feed are not flashy — they matter because they represent the kind of multi-step, multi-tool problem that has been the stubborn gap between test scores and real usefulness. That gap appears to be closing.


