· via Vercel blog
Vercel AI Gateway adds OpenAI Ultrafast tier for GPT-6 Astra and GPT-6.1 Sol
Vercel's AI Gateway can now serve OpenAI's Ultrafast tier on GPT 6 Astra and GPT 6.1 Sol, cutting output latency for interactive apps and coding loops at 6x the standard per-token rate.

Ultrafast arrives on Vercel's AI Gateway
Vercel has added support for OpenAI's Ultrafast service tier on AI Gateway, its routing layer for model requests. According to a changelog post on the Vercel blog, the tier targets workloads where output speed is the main constraint: interactive applications and rapid coding iterations.
At launch, Ultrafast support covers two models: GPT 6 Astra and GPT 6.1 Sol.
How to request it
Developers can request the tier for the model identifiers openai/gpt-6-astra or openai/gpt-6.1-sol through three routes, per Vercel: the AI SDK, the Chat Completions API, and the Responses API.
For workflows that involve frequent tool calls, OpenAI recommends the Responses API over a persistent WebSocket connection, because it reduces the overhead between turns. The changelog points to Ultrafast service-tier examples for persistent connections, including one that runs the AI SDK over WebSocket.
Region support and fallback behaviour
Ultrafast supports US and global processing. Requests pinned to unsupported regions, with the EU named as an example by Vercel, are served at the standard default tier instead.
Standard processing remains the default whenever a request does not specify a service tier, so existing Gateway traffic is unaffected by the change unless it explicitly opts in.
Pricing at six times the standard rate
The speed premium is significant: requests served at Ultrafast are billed at six times the standard per-token rate, according to the Vercel changelog. If a request falls back to another tier, it is billed at the rate of the tier that actually served it, so a request pinned to the EU is not charged the Ultrafast rate for a standard-tier response.
Vercel directs developers to the GPT 6 Astra and GPT 6.1 Sol model pages for current rates, alongside its material on the GPT-6 model family.
Why it matters
For latency-sensitive products such as chat interfaces, autocomplete, and coding agents that loop between model calls and tool execution, generation speed is often the practical bottleneck rather than raw model capability. Exposing a faster service tier through an existing gateway means teams can opt into lower latency without reworking their provider integration: the model identifiers, SDKs, and APIs stay the same, with the tier chosen per request.
The fallback and billing rules keep the trade-offs predictable. A request that cannot be served at Ultrafast for regional reasons runs at standard speed and standard cost, rather than paying a premium for no gain. That predictability matters for teams balancing performance against per-token budgets, since a six-times multiplier only makes sense on the turns where speed genuinely changes the product.
The regional detail is the main caveat. Ultrafast covers US and global processing, so workloads pinned to the EU see no speedup, a consideration for organisations with data-residency requirements that force region pinning.
Finally, the guidance around the Responses API over persistent WebSocket connections signals who this is really aimed at: agentic coding workflows with frequent tool calls, where per-turn overhead can rival the model's own generation time. For those workloads, Vercel's AI Gateway now offers a direct path to OpenAI's fastest tier, with the cost clearly attached.
- #vercel
- #openai
- #cloud
- #ai-gateway
- #llm