deniz.in

Markets

Weather

Loading weather

· via The Verge

Google ships Gemini 3.8 Flash with deeper reasoning at unchanged per-token prices

Google's Gemini 3.8 Flash lands weeks after 3.7 Flash with more reasoning steps and iterative tool calls. Per-token pricing is unchanged, but higher token usage can raise real-world costs.

Google ships Gemini 3.8 Flash with deeper reasoning at unchanged per-token prices

Google has shipped Gemini 3.8 Flash, a rapid follow-up that arrives just weeks after its predecessor, Gemini 3.7 Flash. According to The Verge, Google describes the new model as one that "works harder": it takes more reasoning steps on complex tasks and calls tools iteratively rather than settling for a single pass.

More reasoning, chained tool calls

The headline change is behavioral. The Verge reports that Google frames 3.8 Flash as offering "significant improvements" for software engineering and autonomous agents — the two categories where multi-step reasoning and repeated tool invocation matter most. In practice, the model has been tuned to spend more effort, and more tokens, before it answers.

Per-token prices hold steady, total costs may not

Pricing is unchanged at the introductory level: $0.75 per million input tokens and $3.75 per million output tokens, matching 3.7 Flash. But Google itself cautions that the model may consume more tokens to maximize performance, particularly at higher effort settings.

The benchmarking firm Artificial Analysis put numbers on that effect, as cited by The Verge: despite identical per-token rates, the typical cost of a task with 3.8 Flash runs about 40% higher than with 3.7 Flash, driven by roughly 30% more output tokens per task and additional turns on agentic evaluations. Even so, Artificial Analysis calls it the cheapest model it has measured at this level of intelligence. Developers who want to keep token usage — and bills — down can stay on Gemini 3.7 Flash.

Benchmark results and early impressions

Google says the new model beats its predecessor and other frontier models on the DeepSWE v1.1 software engineering benchmark, finishing ahead of Anthropic's Fable 5 — which received its own upgrade earlier in the week that lowers the price of using cached data. The Verge also reports wins on the Vals Finance Agent V2 benchmark and Harvey's Legal Agent benchmark.

Early reactions were positive. Aigora.ai CEO John Ennis compared the model favorably to Anthropic's lineup, saying it delivers "Opus 5 coding quality but at a fraction of the cost and super fast," and predicted it would be a good fit for generating Remotion videos.

The release also ships with safeguards against misuse in chemical, biological, radiological and nuclear (CBRN) domains as well as offensive cyber operations.

A separate Cyber model for governments

Alongside the general release, Google launched Gemini 3.8 Flash Cyber, a variant restricted to governments and what the company calls trusted partners. Access runs through the new Fairwind Program, whose 650 members include CrowdStrike and the Center for Internet Security, according to The Verge. Members get Flash Cyber plus CodeMender, an agent Google says can autonomously find and fix vulnerabilities in order to protect critical infrastructure, public services and national security.

Availability

Gemini 3.8 Flash is available now to consumers with a Google AI Pro or Ultra subscription, and to developers and enterprise customers through Google's API.

Why it matters

Two things stand out for developers. First, the cadence: a meaningful update landing only weeks after the previous one shows how quickly Google is iterating on its Flash tier, and how little time teams have to stabilize on any given version. Second, the economics of pricing are shifting quietly. Flat per-token rates make 3.8 Flash look free to adopt, but models that reason longer and take more agentic turns spend more tokens per task — Artificial Analysis measured about 40% more here. Anyone budgeting API spend should plan around cost per task, not cost per token. The split release, with a general model alongside a government-only Cyber variant, also signals that vendors are increasingly segmenting frontier capabilities by customer type, a pattern likely to spread across the industry.

  • #google
  • #gemini
  • #llm
  • #ai-agents
  • #api

Related posts