deniz.in

Markets

Weather

Loading weather

· via Vercel blog

Fireworks launches Ember-1 reasoning model on Vercel AI Gateway, touting ~40% fewer tokens than Kimi K3

Fireworks' Ember-1, a reasoning model built on Kimi K3 for coding agents, is now on Vercel AI Gateway as a research preview. Fireworks reports roughly 40% fewer generated tokens than Kimi K3 at comparable quality.

Fireworks launches Ember-1 reasoning model on Vercel AI Gateway, touting ~40% fewer tokens than Kimi K3

Fireworks has put Ember-1, a reasoning model aimed at coding and agentic workloads, on Vercel's AI Gateway. According to a Vercel changelog post published September 28, the model is a research preview built on Kimi K3, and Fireworks' own evaluations found it produces roughly 40% fewer tokens than its base model while holding quality steady.

A shorter reasoning trace, by design

The core pitch is efficiency inside agent loops. A coding agent typically calls its model many times to finish a task, and every call generates reasoning text that must be paid for — and often re-read as context in later steps. As Vercel points out, trimming those intermediate traces can lower output costs and reduce the context an agent carries forward, which helps both with spend and with staying inside a context window.

The 40% figure is a vendor number: Fireworks measured it across its own evaluations, and the results have not been independently verified. It is best treated as a claim to test against your own benchmarks rather than a settled result.

What the model supports

Per the announcement, Ember-1 ships with:

  • A 1M-token context window
  • Text and image input
  • Tool calling
  • Implicit prompt caching
  • Zero Data Retention and No Prompt Training on the Fireworks endpoint

The last two items are the ones teams with strict data policies will check first, since they describe how prompts are handled on the provider side.

Getting started

The model is exposed under the name fireworks/ember-1 across the gateway's APIs and coding agents. Vercel's recommended setup path is two commands: npm i -g vercel@latest to update the CLI, then vercel ai-gateway setup. The setup command scans the machine for coding agents that are already installed, provisions or reuses a gateway API key, and wires each detected agent to route through the gateway; users then pick Ember-1 in their agent's model configuration. A model playground is also available for trying it without touching an existing setup.

AI Gateway itself is a single interface in front of many models, with built-in usage and cost tracking, per-key budgets, and routing rules. The practical effect is that swapping fireworks/ember-1 in or out of an agent is a configuration change rather than an integration project.

A time-boxed preview

Ember-1 ships as a research preview, and the initial run is capped at two weeks. That makes availability explicitly limited, and the model should not be treated as a stable production dependency yet. Teams that want to evaluate it on their own coding tasks have a narrow window to do so.

Why it matters

Reasoning models changed the economics of agents: better answers, paid for with long chains of thought that inflate token bills and fill context windows. A derivative model that keeps quality while emitting fewer tokens attacks that cost directly — and because agent loops compound, with tokens generated early carried into later steps, a 40% cut in output can translate into outsized savings on long-running tasks.

Ember-1 is also a signal of where the model layer is heading: specialized variants of frontier models, tuned for a specific workload such as coding agents and distributed through platforms like Vercel's gateway. For engineering teams, the preview is a cheap experiment — swap the model name, run your evals, and see whether the token savings hold on your own workloads before building any budget plans on top of them.

  • #reasoning-models
  • #coding-agents
  • #vercel
  • #fireworks-ai
  • #ai-gateway

Related posts