· via Vercel blog
Gemini 3.8 Flash arrives on Vercel AI Gateway with 1M token context and 50% discount
Google's Gemini 3.8 Flash is now available through Vercel's AI Gateway, offering a 1M token context window, multimodal input, and a 50% price cut through the end of the year.

Gemini 3.8 Flash lands on Vercel AI Gateway
Google's Gemini 3.8 Flash is now available on Vercel's AI Gateway, and Vercel is discounting the model by 50% through December 31st. According to the Vercel blog, the model can be reached through the gateway using the identifier google/gemini-3.8-flash, alongside the other language models the service already routes.
What the model supports
Vercel lists the following capabilities for Gemini 3.8 Flash:
- a 1 million token context window
- a maximum output of 65,536 tokens
- text, image, PDF and video input, with text output
- tool calling and web search
- thinking enabled by default
Vercel describes the release as an improvement over earlier Flash models in three areas: software engineering, agent-driven work, and multi-step reasoning. The company says those gains arrive at the same speed and price as the previous Flash release.
Wiring up coding agents
For developers using coding agents, Vercel points to its guide on the topic and provides a setup command, vercel ai-gateway coding-agents setup, which connects tools such as Claude Code, OpenCode, Cursor and Pi to the gateway. Once configured, google/gemini-3.8-flash can be selected from within the agent. The model can also be tried directly in Vercel's model playground.
Pricing details
Beyond the time-limited discount, Vercel notes that AI Gateway passes provider pricing through without adding a markup and charges no platform fee on inference. The policy extends to Bring Your Own Key (BYOK) requests, where developers route traffic using their own provider credentials.
Why it matters
Two aspects of the launch stand out. The first is the model itself. Vercel frames Gemini 3.8 Flash as delivering better coding, agent and reasoning performance at an unchanged price and speed — precisely the combination that matters for the agentic and coding-assistant workloads driving much of today's model usage. If the claimed gains hold up in real use, the model becomes an easy candidate to benchmark against incumbents inside existing pipelines.
The second is distribution. Model gateways compete largely on routing convenience and pricing transparency. By mirroring provider pricing with no markup and no inference fee — including for BYOK traffic — Vercel positions its gateway as a neutral switching layer rather than a margin-taking intermediary. A discount running through the end of the year lowers the friction further, giving teams a cheap window to evaluate the new Flash model and, Vercel presumably hopes, to standardize on the gateway while doing so.
- #gemini
- #vercel
- #ai-gateway
- #llm