deniz.in

Markets

Weather

Loading weather

· via Vercel blog

Claude Haiku 5.5 lands on Vercel AI Gateway as Anthropic's fastest standard-speed model

Anthropic's Claude Haiku 5.5, its fastest model at standard speed and the first Haiku with adjustable effort levels, is now available through Vercel's AI Gateway.

Claude Haiku 5.5 lands on Vercel AI Gateway as Anthropic's fastest standard-speed model

Anthropic's Claude Haiku 5.5 is now available through Vercel's AI Gateway, the hosting platform announced in a changelog entry dated October 7, 2026. The model is pitched at workloads where volume and unit cost matter more than frontier capability, and Vercel says it is the quickest Claude model available at standard speed.

What the model is for

According to the Vercel blog, Haiku 5.5 targets frequently executed jobs such as summarization, context compaction, database queries and classification. It can also serve as a subagent next to larger Claude models during coding work, taking on the routine calls while a bigger model handles the complex reasoning.

Vercel suggests treating it as a drop-in replacement wherever Claude Haiku 4.5 is currently used. Given its latency profile at standard speed, the company points to live customer support and browser automation as natural fits.

Tunable effort levels

Haiku 5.5 is the first model in the Haiku line to ship with effort levels. Five tiers are offered — low, medium, high, xhigh and max — and each determines how much the model reasons and how many tokens that reasoning consumes. The feature builds on adaptive thinking: at the low, medium and high settings developers can switch thinking off entirely, while xhigh and max require it to stay on.

In practice, the setting works as a cost dial. A classification job can run at low effort with thinking disabled to keep token usage minimal, while a harder subtask can be pushed to xhigh or max when accuracy justifies the spend. The parameter is set through reasoning in the AI SDK or reasoning_effort in the Chat Completions API.

Safeguards and data handling

The release includes built-in biology and cybersecurity protections, and Vercel notes the model may decline some requests touching those areas. On AI Gateway, Haiku 5.5 supports Zero Data Retention, an option organizations with strict compliance requirements generally want when routing prompts through a third-party proxy.

How to call it

Developers can reach the model through the AI SDK, the Chat Completions, Responses and Anthropic Messages APIs, or via a coding agent connected to the gateway, using the identifier anthropic/claude-haiku-5.5. For tools such as Claude Code, Codex and Cursor, Vercel advises installing the latest CLI, running vercel ai-gateway setup, then selecting the model inside the agent.

Vercel also says the gateway mirrors provider pricing without adding a markup and levies no platform fee on inference, including on bring-your-own-key requests. Beyond routing, AI Gateway provides usage and cost tracking, retries and failover for higher-than-provider uptime, per-key budgets, custom reporting and routing rules.

Why it matters

Most production AI systems are not bottlenecked by frontier intelligence; they are bottlenecked by the cost and latency of thousands of small, repetitive calls — routing a ticket, condensing a conversation, classifying a document. A faster small model with a controllable reasoning budget gives teams a direct lever over that spend, letting them pay for depth only where a task actually demands it.

Bringing effort levels down to the Haiku tier also signals where the industry is heading: reasoning effort is becoming a standard adjustable parameter rather than a premium feature reserved for flagship models. And in multi-agent setups, where a cheap subagent feeds a larger orchestrator, the economics of the small model often decide whether the whole architecture is viable. For teams already routing traffic through Vercel's gateway, Haiku 5.5 reads as a straightforward upgrade path from 4.5 — with the caveat that the new biology and cybersecurity guardrails may refuse requests that earlier versions would have answered.

  • #anthropic
  • #claude
  • #vercel
  • #ai-gateway
  • #llm

Related posts