Claude Haiku 5.5 price cuts, benchmarks reshape model selection
Anthropic released Claude Haiku 5.5 with up to 90% lower token costs and 72.4% computer-use benchmark performance, available now on AWS, Google Cloud, and Azure.

What changed
Anthropic released Claude Haiku 5.5 on October 7, 2026. The company says this is its fastest and most affordable small model to date, designed for high-volume, cost-sensitive tasks like summarization, database queries, classification, and live customer support.
The pricing shift is substantial. According to Anthropic, Haiku 5.5 costs about 75 percent less than Haiku 4.5 on average. For requests with prompts up to 100,000 tokens, which Anthropic says account for roughly 90 percent of all previous Haiku requests, prices drop by up to 90 percent. Prompts longer than 100,000 tokens cost five times as much.
The company notes that Haiku 5.5 uses an updated tokenizer that consumes slightly more tokens per task than its predecessor. Anthropic points out that the same thing happened with Opus 4.x models, where token usage jumped about 30 percent from the tokenizer change alone. This means real-world savings will likely be smaller than per-token price comparisons suggest.
Haiku 5.5 is the first Haiku-class model with adjustable reasoning levels, letting users balance cost against quality. The model is available now across Amazon Web Services, Google Cloud, and Microsoft Azure.
As part of the same release, Anthropic cut cache read costs for Sonnet 5.5 by 50 percent, from $0.20 to $0.10 per million tokens. The company also introduced monthly API credits for Claude Max and Team subscribers, ranging from $100 to $500 per month.
Why it matters
Claude Haiku 5.5’s performance improvements at drastically lower costs reshape when development teams should use smaller models instead of larger ones. The benchmark results matter most here.
Haiku 5.5 scores 72.4 percent on OSWorld 2.1, a computer use benchmark, up from 15.7 percent for Haiku 4.5. On the agentic coding benchmark Terminal-Bench 4.0, it reaches 39.2 percent while Haiku 4.5 scored zero. For knowledge work on GDPval-AA v2.1, it scores 1,620 compared to 735 for Haiku 4.5, more than double. On Humanity’s Last Exam, it hits 45.9 percent without tools and 57.4 percent with tools, up from 10.2 and 18.7 percent respectively.
These gains open up use cases that previously required larger, more expensive models. Computer use tasks, which burn through large amounts of tokens, become practical with a cheap, fast model. Agentic work that previously needed Sonnet or Opus can now run on Haiku 5.5 as a subagent, cutting costs significantly while maintaining acceptable performance.
The pricing changes also matter for infrastructure budgets. Organizations running millions of API calls for customer support, document analysis, or data classification can reduce expenses substantially. Asana reported over 30 percent latency reduction and up to 2.5x faster inference per agent turn compared to their previous setup. HubSpot achieved 92.8 percent on its CRM evaluation suite using Haiku 5.5, the highest score they have seen.
The Sonnet 5.5 cache read price cut, reducing agentic work costs by about 20 percent according to Anthropic, signals an intensifying price war. The company explicitly compares Haiku 5.5 to OpenAI’s GPT-6 Luna budget model, and says Haiku 5.5 leads across every tested category.
What to test
Before committing workloads to Haiku 5.5, teams should verify several vendor claims:
Benchmark applicability. The benchmark improvements are real according to Anthropic’s published results, but check whether they reflect your actual use cases. Computer use tasks, for instance, may not match your production environment exactly.
Real-world token consumption. Test actual prompts and tasks to measure token usage with the new tokenizer. Anthropic notes that token consumption increased about 30 percent when Opus 4.x moved to the new tokenizer. Calculate whether the per-token price savings still apply after accounting for higher token counts in your workloads.
Cost per task, not per token. Compare total cost per completed task between Haiku 5.5 and your current model. Latency matters too. A 30 percent reduction in latency, as Asana reported, may be significant for real-time applications but irrelevant for batch processing. AlphaSense found Haiku 5.5 was a statistically significant improvement on their document query work, with quality scores of 0.84 versus 0.76 for Haiku 4.5.
Prompt length variations. The 90 percent price cut only applies to prompts up to 100,000 tokens. Prompts longer than that cost five times as much. Measure how many of your requests exceed this threshold. If a significant portion do, calculate blended costs accordingly.
Adjustable reasoning quality. Test Haiku 5.5 with different effort settings to find the cost-quality balance that works for your tasks. Anthropic says the model works best for narrowly scoped tasks like summarization, compaction, or sub-agent work but remains less suitable for complex agentic coding compared to Sonnet 5.5 and Opus 5.5.
Comparison against alternatives. While Anthropic shows Haiku 5.5 beating GPT-6 Luna on tested benchmarks, run your own evaluations against models you currently use. Third-party benchmarks may differ from Anthropic’s results.
Safety and compliance. Verify that Haiku 5.5’s cybersecurity safeguards meet your requirements. Anthropic blocks penetration testing but allows a broader range of defensive tasks than Sonnet 5.5 does. If your work touches restricted areas, check the Life Sciences Verification Program and Cyber Verification Program requirements.
The conclusion
Claude Haiku 5.5 represents a meaningful shift in small-model economics. The combination of 72.4 percent computer use performance and up to 90 percent lower token costs on typical requests makes it a credible replacement for Haiku 4.5 and potentially for more expensive models on many tasks.
The caveat is that token consumption increased with the new tokenizer. Real-world savings will be lower than the headline figures suggest, especially for long prompts. Teams need to test their specific workloads before migrating. The adjustable reasoning levels help here, allowing cost-quality experimentation.
What stands out is the competitive context. Anthropic’s price cuts and new API credits follow OpenAI’s GPT-6.1 release, confirming that AI model pricing remains contested. Both companies are competing hard on small-model economics. This benefits customers, but it also suggests prices could shift again soon.
Watch whether Haiku 5.5’s computer use capabilities (72.4 percent on OSWorld 2.1) prove durable in production environments. Early feedback from Asana, HubSpot, Box, and others suggests the model delivers on those benchmarks, but those are vendor testimonials rather than independent verification. Also monitor whether other companies match Anthropic’s price cuts or push performance further.
For teams currently using Haiku 4.5 or larger models for high-volume, low-complexity tasks, migrating to Haiku 5.5 is worth testing now. For those on other providers’ small models, compare directly on your own data. The economics have shifted enough that switching is plausible for many workloads.
AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure
Related stories

Claude Dynamic Workflows: 1,000 Parallel Agents Now Available
Anthropic added dynamic workflows to Claude Managed Agents, enabling up to 1,000 AI agents to run in parallel per execution through managed agent infrastructure.

Microsoft releases Decision-1 routing and classification model
Microsoft releases Decision-1, a Qwen3.5-9B decision-scoring model for routing and classification tasks, available in Foundry and OpenRouter.

Nace AI Open-Sources Drex 1.5 Decision Model for Option Scoring
Nace AI open-sources Drex 1.5, a 9B decision model that scores multiple options in one forward pass, ranking first under 10B parameters on Decision Index.