Anthropic Unveils Claude Haiku 5.5 with Steep Rate Cuts and Benchmark Upgrades
Anthropic has introduced Claude Haiku 5.5, delivering lower per-token pricing, adjustable reasoning controls, and significant benchmark gains in computer use and agentic workflows.

Anthropic has officially launched Claude Haiku 5.5, marking its fastest and lowest-cost lightweight model to date. Aimed directly at high-volume, cost-conscious workloads, the new release is tailored for use cases such as content summarization, classification, database queries, and real-time customer service. Alongside this launch, Anthropic has also lowered cache read rates for its mid-tier model, Sonnet 5.5, highlighting the continued competitive price adjustments occurring across the artificial intelligence sector.
On standard requests, Haiku 5.5 costs approximately 75 percent less on average than Haiku 4.5. For prompts containing up to 100,000 tokens—a threshold Anthropic notes covers roughly 90 percent of historic Haiku traffic—prices are reduced by up to 90 percent. Under this sub-100,000 token bracket, input tokens cost $0.10 per million and output tokens cost $0.50 per million, down from $1.00 and $5.00 on Haiku 4.5. Cache reads drop from $0.10 to $0.01 per million, while cache writes fall from $1.25 to $0.125. However, longer contexts exceeding 100,000 tokens incur a fivefold rate increase, priced at $0.50 per million input tokens, $2.50 per million output tokens, $0.05 for cache reads, and $0.625 for cache writes.
Significant benchmark gains and computer use
Performance evaluations show substantial gains across diverse tasks compared to the prior generation. On the knowledge benchmark GDPval-AA v2.1, Haiku 5.5 registered a score of 1,620, more than doubling the 735 scored by Haiku 4.5 and surpassing OpenAI's budget model, GPT-6 Luna, which scored 1,437. Similarly, on AA-Briefcase v1.1, Haiku 5.5 scored 1,578 against Haiku 4.5's 614.
Reasoning and coding metrics show similar movement. On Humanity's Last Exam, Haiku 5.5 reached 45.9 percent without tools and 57.4 percent with tools, compared to 10.2 percent and 18.7 percent for Haiku 4.5. On the agentic coding test Terminal-Bench 4.0, Haiku 5.5 achieved 39.2 percent after Haiku 4.5 failed to score, also outperforming GPT-6 Luna's 16.4 percent. On FrontierCode 1.1, Haiku 5.5 marked 46.4 percent against GPT-6 Luna's 42.4 percent.
The most notable leap occurred in autonomous interface tasks. On the OSWorld 2.1 offline benchmark, Haiku 5.5 scored 72.4 percent, up from 15.7 percent on Haiku 4.5 and ahead of GPT-6 Luna's 48.9 percent. Because operating a desktop environment consumes high volumes of tokens rapidly, a fast, low-cost model is structurally well suited for such workflows. Even so, benchmark references show Sonnet 5.5 remains noticeably higher across categories, including 83.9 percent on OSWorld 2.1 and 70.6 percent on Terminal-Bench 4.0.
Technical adjustments and system capabilities
Haiku 5.5 is the first model in the Haiku family to introduce adjustable reasoning levels, giving engineers direct control to trade execution speed and cost against output depth. Anthropic recommends reserving Haiku 5.5 for defined operations like data compaction or auxiliary agent tasks, while directing demanding software engineering toward Sonnet 5.5 or Opus 5.5.
Developers evaluating total expenditure must also account for a newly updated tokenizer. As seen in previous model revisions like Opus 4.x—where tokenizer adjustments expanded token consumption by roughly 30 percent—Haiku 5.5 consumes slightly more tokens per task than Haiku 4.5. As a result, net operational savings may be somewhat lower than the per-token price reductions indicate.
Safety guardrails have also been adjusted. While penetration testing remains restricted, Haiku 5.5 permits a wider range of defensive security tasks than Sonnet 5.5 due to differences in overall capability. Qualified organizations can also apply for Anthropic's verification programs in cybersecurity and life sciences.
What it means for developers
For engineering teams managing large-scale infrastructure, the arrival of Haiku 5.5 offers a much cheaper option for high-frequency operations and sub-agent architectures. Because the model handles routine logic and computer control at a fraction of larger models' cost, developers can offload simple agent steps from pricier tiers without facing severe performance bottlenecks.
Access is open immediately across Amazon Web Services, Google Cloud, and Microsoft Azure, alongside official Python and TypeScript SDKs that include beta support for browser and computer use. Developers exploring multi-model pipelines can also try top AI models cheaply through one API at https://apixoai.online.
Concurrently, Anthropic halved cache read costs on Sonnet 5.5 from $0.20 to $0.10 per million tokens, cutting overall expenses for agentic workloads by roughly 20 percent. To incentivize implementation, Anthropic is distributing monthly API credits ranging from $100 for Max-5x users to $200 for Max-20x users and up to $500 for Team subscribers to experiment with custom tooling and agents.
Source: Claude Haiku 5.5 arrives with massive price cuts proving the AI pricing arms race is far from over — The Decoder. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

