· via Vercel blog
Vercel AI Gateway escalates low-confidence AI decisions to fallback models
Vercel's AI Gateway can now reroute decision requests to a fallback model when the primary answer falls below a confidence threshold, adding an output-quality safety net for production AI.

What changed
Vercel has added confidence-based decision fallbacks to AI Gateway, its routing layer for AI model traffic. According to the Vercel blog, the gateway can now hand a decision request off to a fallback model whenever the answer coming back from the primary model fails a confidence check the developer defines — for instance, a confidence score under 0.6 on a particular question.
Until now, fallbacks in the gateway worked as error safety nets: list several models, and if one fails outright, traffic moves to the next entry. Vercel says that behaviour is untouched — plain model names in the models list still catch hard errors. The new conditional objects target a different problem: responses that arrive successfully but are too uncertain to act on.
How the conditions work
Confidence conditions cover Choice and Score questions, while Boolean questions escalate on a probability range instead. Conditions can also be combined, so a single escalation can depend on more than one signal before the request is rerouted.
Two levels of scoping are available. Name a specific question in the condition and only that question is checked. Leave the question field out and the condition applies to every question of the matching type — Vercel notes that a bare 0.6 threshold escalates as soon as any Choice or Score answer comes in under it.
Configuration
The feature is configured per request inside providerOptions.gateway.models. Each entry pairs a fallback model with a "when" object describing the condition:
{ "gateway": { "models": [ { "model": "openai/gpt-6-astra", "when": { "question": "intent", "confidenceBelow": 0.6 } } ] } }
Requests that leave the conditional object out behave exactly as they did before, which makes the feature opt-in and safe to roll out gradually.
Costs and caveats
There is a billing consequence worth flagging. When a threshold trips, the gateway runs a second decision on the fallback model, and both stages are billed. Teams running high volumes will therefore want to tune thresholds carefully: a cutoff set too high turns the feature into a near-constant double spend, while one set too low rarely fires at all. Vercel also lists decision fallbacks as a beta feature.
Why it matters
Outright errors were already handled by model lists. What production AI pipelines have lacked is a routing policy based on the quality of the output itself. Low-confidence answers are a quiet failure mode: the system responds, the pipeline continues, and the weak decision only surfaces downstream. By letting developers escalate just the uncertain slice of traffic to a stronger or differently tuned model, Vercel offers a reliability lever that requires no rearchitecting — only one extra object in the request. For decision-heavy workloads such as intent detection, classification or moderation, that can translate directly into fewer actions taken on shaky model output.
- #vercel
- #ai-gateway
- #llm
- #fallback-routing
- #developer-tools