· via dev.to (home feed)
Gemini 3.8 Flash ships with unchanged pricing, doubled 2027 rates and a gated Cyber model
Google's third Flash release in three months holds introductory pricing until the end of 2026, but added reasoning steps can inflate output-token bills. A defense-only Cyber variant ships under a restricted access program.

Third Flash release in three months
Google shipped Gemini 3.8 Flash on September 2, 2026, three weeks after Gemini 3.7 Flash and the third Flash model in as many months, according to a Proje Defteri write-up on dev.to. The post dates 3.6 Flash to July 21 and 3.7 Flash to August 13. Google reportedly pitches the new model as its strongest reasoning and coding release while keeping the speed and cost profile of 3.7, in an announcement signed by Tulsee Doshi, senior director of product management, and Raluca Ada Popa, Gemini security lead at Google DeepMind.
The published spec sheet lists model ID gemini-3.8-flash, a one-million-token context window, a 64,000-token output cap, multimodal input covering text, image, video, audio and PDF with text-only output, plus function calling, search grounding and computer use. Weights stay closed. The knowledge cutoff is March 2026 for most domains but January 2025 for some, which the post flags as a hallucination risk for questions about recently released libraries unless grounding is enabled.
More reasoning steps, bigger invoices
The central behavioral change is that the model works harder: extra reasoning steps on complex tasks, iterative tool calls instead of one-shot calls, and less willingness to settle for a first answer. That cuts both ways for API users. Thinking tokens are billed at the output rate, five times the input rate, so identical prompts can return noticeably more billable tokens. Artificial Analysis, cited by the post, rates the model as very verbose and logged 120 million output tokens across its evaluation suite. Google recommends lower effort levels for efficiency-first workloads and says 3.7 Flash remains fully supported for those jobs.
On speed, Artificial Analysis measures 304.6 output tokens per second, the fastest in its index, but 13.39 seconds to first token, a delay users feel in chat interfaces though not in batch pipelines.
Benchmarks split between specialist and long-horizon work
The benchmark compilation in the post, drawn from Google, Artificial Analysis and an OfficeChai roundup, shows a divided picture against Claude Opus 5 and GPT-5.6 Sol. 3.8 Flash leads the table on specialist agent work: 61.4 percent on Vals Finance Agent v2, 10.0 percent on Harvey's legal agent benchmark against 6.7 for Opus 5 and 2.5 for GPT-5.6 Sol, 54.9 percent on HLE-Verified, 86.2 percent on CharXiv chart reasoning, and 87.8 percent on agentic LVBench long-video understanding versus 75.4 for Opus 5.
Long-horizon autonomous work is where it trails: 19.1 percent on Terminal-bench 4.0 against Opus 5's 51.8, 59.0 percent on OSWorld-2.0 computer use against 75.4, 71.0 percent on DeepSWE v1.1 behind both rivals, and a GDPVal-AA v2 Elo of 1545 versus 1824. Gains over its own predecessor are real, with Terminal-bench 4.0 up from 11.2 percent and the hard BioMysteryBench set up from 43.5 to 56.5 percent, and the AA Intelligence Index moving from 56 to 59, 16th out of 195 models. The post notes Opus 5 costs $5 input and $25 output per million tokens, roughly 6.7 times Flash on both axes, so routing critical steps to the larger model and volume work to Flash remains the economical architecture.
Pricing holds until it doesn't
Introductory API pricing is unchanged at $0.75 input and $3.75 output per million tokens, but it expires December 31, 2026, doubling to $1.50 and $7.50 from January 1, 2027. The post points out this is the same tariff and expiry as 3.7 Flash, and advises budgeting for both the New Year increase and token inflation from the added verbosity.
Flash Cyber ships behind a new access gate
A second model arrived the same day: Gemini 3.8 Flash Cyber, a defense-only successor to 3.5 Flash Cyber tuned for autonomous vulnerability discovery. Google claims over 70 percent success on real-world vulnerability discovery across 20 programming languages, 47.2 percent pass@1 on CWE-Bench patching, improved Gray Swan prompt-injection robustness, and wins over its predecessor and larger frontier models on CyberGym. Field reports cited include the Chrome security team seeing 2.6 times more correct patches than leading commercial models, Wiz reporting 7.5 to 9.7 percent higher recall at 2.3 to 5.2 times lower cost, and Google Cloud's vulnerability research team crediting it with finding a critical flaw in under two hours. It is not callable with a standard API key: access is limited to trusted defenders through the new Fairwind Program, which the post says includes government authorities.
Why it matters
For API users the release is less about a price move and more about two deadlines and one hidden cost. The headline rate is unchanged, but it doubles on January 1, 2027, and heavier reasoning can quietly raise output-token spend before that date arrives. The benchmark split also shapes architecture: 3.8 Flash is the strongest model in its price class for finance, legal, chart and long-video agent work, while long-horizon terminal and computer-use tasks still justify paying roughly seven times more for Opus 5. Meanwhile, Flash Cyber shows Google segmenting its most security-capable models away from general API access, a pattern worth watching as agentic security tooling matures.
- #gemini
- #llm
- #api-pricing
- #benchmarks
- #cybersecurity