Technology

Claude Haiku 5.5: Anthropic's Cheapest, Fastest Model for AI Agents

Martin HollowayPublished 16m ago4 min readBased on 5 sources
Reading level
Claude Haiku 5.5: Anthropic's Cheapest, Fastest Model for AI Agents
source:anthropic.com

Anthropic introduced Claude Haiku 5.5 on Oct. 7, 2026, completing the initial rollout of its Claude 5.5 family with a small model it describes as its cheapest, fastest, and most capable to date. Anthropic

On cost, Anthropic states Haiku 5.5 costs around 75% less to run on average than Haiku 4.5. It also describes Haiku 5.5 as its fastest model to date.

The model is built for high-volume, cost-sensitive inference, meaning live use where the model answers many requests. Anthropic lists summaries, compactions, database queries, and classification requests as the design center. In those jobs, per-call delay and per-token price drive total system cost, especially in agentic loops that issue hundreds or thousands of smaller calls.

Haiku 5.5 is also the first Haiku-class model with an adjustable effort setting to trade cost against intelligence. The control will be familiar to users of the larger Opus and Sonnet tiers. Low effort fits bulk triage and compaction. Higher effort fits cases where a small-model call sits on a critical path.

Anthropic positions Haiku 5.5 as a complement to larger models, not a replacement. It states the model pairs with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. That setup is now common in multi-agent systems. A frontier model plans and reviews. A fleet of smaller models handles retrieval, edits, test runs, and summarization.

Measured capability

On knowledge-heavy evaluations, Anthropic reports Haiku 5.5 scored 1620 on GDPval-AA v2.1 and 1578 on AA-Briefcase v1.1. On Humanity's Last Exam, it reports 45.9% without tools and 57.4% with tools.

The wider point for agent builders here is the gap between tool-off and tool-on scores. It shows how much of the model's useful ability depends on retrieval, code execution, and structured tool access.

On computer use and visual reasoning, Anthropic reports 72.4% on the OSWorld 2.1 offline subset and 46.4% without tools on Chartography. Anthropic

On code and terminal work, Anthropic reports 46.4% on FrontierCode 1.1 Main and 39.2% on Terminal-Bench 4.0. These benchmarks combine long-horizon planning, shell interaction, and repository-scale edits.

In practical terms here, scores in that range point to a supporting role rather than a lead role on frontier software engineering. They point to a model that can take on bounded subtasks reliably enough to be dispatched at scale.

Early deployment data points in the same direction. Asana reported over 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn in early testing of Haiku 5.5.

The practical consequence here for multi-turn agents is that time per turn adds up quickly across turns. A 2.5x gain changes interaction design and timeout budgets.

Pricing and rollout

Anthropic paired the model release with two platform pricing changes. It is halving the price of Claude Sonnet 5.5's cache reads. After that cut, Anthropic states Sonnet 5.5 now runs around 20% cheaper on most agentic work. Cache pricing matters for long-context agents, where prompt caching, repeated system prompts, and shared prefixes make read pricing as important as base input pricing.

Anthropic is also introducing a new monthly API credit for Claude Max and Team subscribers to support building agents and applications on the Claude Platform. Details on credit amounts were not included in the verified announcement, but the structure is clear. Subscription spend converts directly into platform experimentation.

The release completes a sequence Anthropic had telegraphed through September. Anthropic dated its Opus 5.5 announcement to Sep. 22, 2026 and its Sonnet 5.5 announcement to Sep. 28, 2026, with Haiku 5.5 dated Oct. 7, 2026. Earlier statements said Sonnet 5.5 and Haiku 5.5 would follow Opus 5.5 in the coming weeks, with many of the same improvements to performance, efficiency, and safety. Anthropic

The broader family trend is efficiency. Anthropic describes Sonnet 5.5 as running 30% faster and costing up to 30% less for most work than Sonnet 5. It states Opus 5.5 performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. Separately, Sonnet 5.5 is priced at $2 per million input tokens and $10 per million output tokens. Reuters

The broader context here is straightforward for teams budgeting agent infrastructure. Frontier capability still gets attention, but deployment cost now sets architecture. When a small model cuts operating cost by three quarters while cutting latency by a third or more, engineers delegate more aggressively. They add verification steps, parallel candidates, and compaction passes that were previously too expensive to run.

In my view, the adjustable effort setting may prove as consequential as the raw price cut. It lets one model serve two roles. Bulk classifier in one call, careful subagent in the next, without redeploying endpoints or managing separate version pins. That simplifies routing logic and observability, which are often the real bottlenecks in production agent systems.

I have seen this cycle before at home, watching my own children move from treating every search query as precious to treating generation as disposable, once speed and price crossed a threshold. Enterprise developers make the same shift. Cheap, fast subagents invite looser invocation and tighter feedback loops. The risk is sprawl. The opportunity is systems that check their own work as a matter of course. Haiku 5.5 looks built for that second outcome.