· via Hacker News – Front Page (hnrss.org)
Anthropic launches Claude Haiku 5.5, a small model priced about 75% below Haiku 4.5
Anthropic's Claude Haiku 5.5 targets high-volume, cost-sensitive workloads, runs about 75% cheaper than Haiku 4.5, and arrives with a Sonnet 5.5 cache-read price cut and new API credits for subscribers.

Anthropic has released Claude Haiku 5.5, which the company calls its cheapest and fastest small model yet, and its most capable. The announcement on Anthropic's site, currently on the Hacker News front page, is aimed at developers running high-volume, cost-sensitive workloads where per-token price and latency matter more than peak reasoning ability.
A workhorse model for narrow jobs
Haiku 5.5 is positioned as a workhorse rather than a frontier model. According to Anthropic, it reliably handles quick, repetitive work such as summarization, context compaction, database queries and classification, and it can act as a subagent alongside the larger Opus 5.5 and Sonnet 5.5 in coding setups, handling narrow steps that do not justify a bigger model's price tag. Because it is Anthropic's fastest model to date, the company also points to latency-sensitive uses like live customer support and browser automation.
Roughly 75% cheaper to run
Anthropic says Haiku 5.5 costs on average around 75% less to run than Haiku 4.5. Pricing is tiered by prompt length. For prompts up to 100,000 tokens — which Anthropic says make up about 90% of requests to the previous Haiku model — input costs $0.10 per million tokens and output $0.50 per million, with cache reads at $0.01 and cache writes at $0.125. Prompts over 100k tokens are billed at $0.50 per million input, $2.50 per million output, $0.05 for cache reads and $0.625 for cache writes. By comparison, Haiku 4.5 charged $1.00 per million input tokens and $5.00 per million output tokens.
Benchmark gains, with caveats
Anthropic's published benchmarks show a large jump over Haiku 4.5. On OSWorld 2.1 (offline subset), which tests computer use, Haiku 5.5 scores 72.4% against Haiku 4.5's 15.7%; Sonnet 5.5 reaches 83.9% and a model listed in Anthropic's table as GPT-6 Luna scores 48.9%. On Terminal-Bench 4.0, an agentic coding benchmark, Haiku 5.5 manages 39.2% while Haiku 4.5 scored 0.0% and Sonnet 5.5 reaches 70.6%. On Humanity's Last Exam without tools, Haiku 5.5 posts 45.9% versus Haiku 4.5's 10.2%. Anthropic itself notes the gap to its larger models: Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding, while Haiku 5.5 fits narrowly scoped jobs that previously would have been too expensive to hand to a Claude model.
Early customer feedback quoted in the announcement is consistent with the speed claims. Asana staff software engineer Aaron Vinh reported a reduction of more than 30% in task-completion latency and up to 2.5x faster inference per agent turn compared with the model Asana uses today.
First Haiku with adjustable effort
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, so developers can tune the trade-off between cost and intelligence on a per-task basis, as they already can with Anthropic's larger models.
Wider pricing moves across the lineup
Alongside the launch, Anthropic halved cache-read pricing for Claude Sonnet 5.5, from $0.20 to $0.10 per million tokens. Because cache reads account for a large share of token consumption in agent loops, Anthropic estimates this makes Sonnet 5.5 roughly 20% cheaper on most agentic work.
The company is also rolling out a monthly API credit for Claude Max and Team subscribers: $100 per month for Max 5x users, $200 for Max 20x users, and up to $500 pooled across users on Team plans, spendable on any model. In addition, the Claude Python and TypeScript SDKs are gaining beta support for computer use and browser use, tasks Anthropic says Haiku 5.5 handles especially well given its speed and price.
Safety posture
On safety, Anthropic reports fewer instances of misaligned behavior and a lower willingness to cooperate with misuse compared with Haiku 4.5. Haiku 5.5's cybersecurity safeguards are more restrictive than Haiku 4.5's but somewhat looser than those on other recent models: they allow a wider range of defensive security tasks than Sonnet 5.5's safeguards while still blocking penetration testing and similar offensive techniques. Its biology safeguards match those applied to Sonnet 5, Sonnet 5.5 and Opus 5, permitting research questions while restricting requests judged likely to cause harm.
Availability
Claude Haiku 5.5 is available now on the Claude Platform under the model identifier claude-haiku-5-5, and through Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic has published a migration guide for developers moving over.
Why it matters
Small models are where the economics of applied AI get decided. Most real-world agent pipelines spend the bulk of their tokens on routine steps — routing, summarizing, compressing context — and a roughly 75% price cut on those steps changes which automated systems are financially viable. The tiered, prompt-length-based pricing is also notable: it prices short requests far below long ones, acknowledging where the volume actually sits. For rivals building competing small models, Anthropic has set an aggressive reference point, and the subscriber API credits push customers toward building on the Claude platform rather than only using its apps.
- #anthropic
- #claude
- #llm
- #api-pricing
- #ai-agents