· via Vercel blog
Open-weight models reach 56% of production token volume on Vercel's AI Gateway
Vercel's September AI Gateway index shows open-weight models carried 56% of production tokens in August — a first — while average token prices fell 23.2% and Anthropic held 64% of all spend.

Open-weight models pass half of production tokens
Open-weight models carried 56% of all tokens routed through Vercel's AI Gateway in August 2026, the first month they accounted for a majority of production volume on the service. According to Vercel's September Production Index, published on the company's blog, open-weight models processed fewer than one in ten gateway tokens in December 2025 and 13% in April, then rose every month through August. The dollar picture is different: open-weight models accounted for 56% of August's tokens but only 14% of spend, leaving closed-weight frontier models with most of the money.
Price per token keeps sliding
The average price per token across the gateway fell 23.2% in August, the third consecutive monthly drop and the steepest since April, according to the report. Vercel estimates the average token now costs less than half of what it did five months ago. Among teams that ran more than ten million tokens in both July and August, the median team paid 7.6% less per token, more than double July's 2.9% decline. Vercel ties the falling averages partly to the shift toward open-weight models, noting that teams can now run more inference on the same budget and reserve frontier models for tasks that justify the premium.
Teams step down from Fable 5 to Opus 5
Fable 5, which Vercel describes as Anthropic's most capable model, fell from 13.2% of gateway spend in July to 4.9% in August. Over the same period, Opus 5, priced at roughly half of Fable's per-token rate, rose to 22.5% of spend. Nine in ten teams that ran Fable cut their usage, and more of them moved workloads to Opus 5 than to any other model. Because those workloads stayed inside Anthropic's lineup, the lab kept 64% of all gateway spend in August, and Vercel notes Anthropic has taken at least 61 cents of every gateway dollar every month since December.
The July context matters here: a US export control on Fable 5 was lifted and access restored on July 1, pushing Fable's spend share to 13.2%, before Opus 5 came online at the end of that month.
Gemini 3 Flash loses most of its token share
Google fared worse. Gemini 3 Flash has lost 95% of its share of gateway tokens since May, and more than three-quarters of that volume went to models from other labs; Google's own full-size Flash successors picked up under a tenth of it. About half of the departing volume went to cheaper models, led by GPT-5.6 Luna, and most of the other half to higher-priced models, led by Claude Opus 5 and Sonnet 5. Google's share of gateway token volume fell from 30% to 5%, with Gemini 3 Flash accounting for 22 of the 25 percentage points lost.
Vercel's read is that customer loyalty follows the model profile rather than the lab. Z.ai's GLM-5.3-Flash, for instance, overtook GLM-5.2 one day after appearing on the gateway and was processing two-thirds of Z.ai's tokens by August 31.
Astra takes a fast lead over Fable 5.1
OpenAI launched GPT-6 Astra on the gateway on September 3 at the same price as Anthropic's Fable 5.1, which arrived two days earlier. Within two days, Astra accounted for one in every three dollars spent on OpenAI models through the gateway, and its share has hovered between 28% and 39% since. Over each model's first twelve days, Astra took 7.7% of all gateway spend, more than twice Fable 5.1's 3.7%, and was used by twice as many teams. From September 4 to 16, Astra and Sol processed 27% of OpenAI's tokens but 71% of its spending, while Luna handled more than eight times Astra's token volume even as Astra accounted for more than four times Luna's spend.
Elsewhere in the data
Google's Nano Banana took the lead in image spend for the first time, at 50% to GPT Image's 44%, though GPT Image stayed ahead on images generated, 46% to 39%. In video, Veo rose to second in spend at 20%, behind Seedance in both videos generated and video spend, while xAI's Grok Imagine saw its share of videos generated more than halve since June, from 42% to 19%.
Why it matters
Vercel says AI Gateway routes tens of trillions of tokens each month between production applications and AI labs, which makes the index a rare measure of what teams actually deploy rather than what they benchmark. The crossover to a majority of token volume running on open-weight models, combined with a third straight monthly price decline, signals that open models now carry real production workloads, not just experimentation, while capturing comparatively little revenue. For labs, the August data suggests pricing tiers and continuity matter as much as peak capability: Anthropic lost two-thirds of its flagship's spend share yet kept its customers inside its lineup, while Google's successors, which Vercel says offered no relative advantage on capability or price, sent most departing volume to competitors. The figures cover traffic through one gateway, and Vercel estimates spend using labs' published list prices, so actual bills may differ.
- #open-weight-models
- #ai-infrastructure
- #llm
- #vercel
- #token-pricing