deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Google launches Gemini 3.8 Flash and a defenders-only Flash Cyber model

Google's third Flash release in six weeks pairs an agentic coding model with a cybersecurity variant that finds and patches vulnerabilities, available only to vetted defenders under the new Fairwind Program.

Google launches Gemini 3.8 Flash and a defenders-only Flash Cyber model

Google's third Flash release in six weeks

Google has introduced two new models: Gemini 3.8 Flash, a general-purpose model aimed at long-horizon coding and autonomous agents, and Gemini 3.8 Flash Cyber, a cybersecurity variant built for vulnerability discovery and automated patching. According to Google's announcement on 2 September 2026, this is the company's third Flash release in six weeks, arriving three weeks after 3.7 Flash.

Both variants share the same underlying model. Google says that shared core was trained heavily on cybersecurity problems and refined through long-running agentic loops that recursively evaluate and improve the models.

Pricing for the general model stays at 3.7 Flash's introductory level: $0.75 per million input tokens and $3.75 per million output tokens. Flash Cyber is not on open sale — it is restricted to vetted defender organisations through a new access scheme Google calls the Fairwind Program.

A model that trades tokens for results

The central design choice for 3.8 Flash, per Google, is persistence: on complex tasks it runs additional reasoning steps and calls tools iteratively, which can consume more tokens at higher effort levels. Developers who care most about efficiency can dial effort down or keep using 3.7 Flash, which remains supported for workloads where efficiency matters most.

On benchmarks, Google claims 3.8 Flash outperforms most larger frontier models on DeepSWE v1.1, a long-horizon software engineering test, at a fraction of their cost. The company also reports leads on Vals Finance Agent V2 and Harvey's Legal Agent Benchmark, plus a 54.9% score on HLE-Verified for multi-step reasoning across STEM, humanities and professional domains. These are Google's own reported figures.

Cyber variant focused on patching, not exploitation

Google says it deliberately prioritised vulnerability fixing over offensive capabilities such as exploitation. On CyberGym, an industry benchmark for finding vulnerabilities in C and C++ code, Flash Cyber reportedly surpasses both the earlier 3.5 Flash Cyber and significantly larger frontier models. On an internal benchmark spanning codebases written in 20 programming languages, Google reports a success rate above 70%.

On CWE-Bench, an external patching benchmark run by Collinear, Flash Cyber achieved a pass@1 of 47.2% versus 47.8% for a leading frontier model, at what Google describes as significantly lower cost.

Google also shared deployment results from inside and outside the company. The Chrome Security team found Flash Cyber produced 2.6 times more correct vulnerability patches than much larger commercial models. Security firm Wiz reported 7.5 to 9.7 percentage points higher recall on its internal penetration-testing benchmark at 2.3 to 5.2 times lower cost. Google's Cloud Vulnerability Research team used the model to find a critical foundational vulnerability in under two hours, work that ordinarily takes months to research.

What the model card adds

A Google DeepMind model card published the same day fills in technical and safety detail. Gemini 3.8 Flash is based on 3.7 Flash, accepts text, images, audio and video, has a context window of up to 1 million tokens with a 64K token output limit, and keeps the family's customisable effort levels for balancing quality, cost and latency.

The card lists limitations: general foundation-model issues such as hallucination, occasional slowness or timeouts, and a knowledge cutoff of March 2026, with some domains effectively limited to January 2025 in line with the wider Gemini 3 family.

On safety, the card tempers the launch messaging somewhat. Google's announcement highlights a significant leap in prompt-injection robustness as measured by Gray Swan, but the model card notes multilingual safety regressed by 5.4 percentage points against 3.7 Flash, with text-to-text safety essentially flat. A Frontier Safety assessment under Google's April 2026 framework concluded 3.8 Flash shows no meaningful new capabilities over its predecessor and is unlikely to reach any tracked or critical capability levels. Human red teaming found no egregious concerns, according to the card.

Flash Cyber ships with looser cybersecurity mitigations than the standard model, which is precisely why Google limits it to defenders it has vetted.

Why it matters

Three points stand out. First, cadence: three Flash releases in six weeks signals that Google now iterates its mid-tier models faster than many rivals refresh flagships, treating them as the real workhorses of agent deployment. Second, price-performance: near-frontier results on agentic, coding and security benchmarks at $0.75 per million input tokens pressures competitors, especially for long-running agents where token costs compound. Third, governance: a defenders-only security model, paired with published benchmark and safety data, sketches a template for how dual-use AI capabilities might be distributed without wide release.

  • #google
  • #gemini
  • #ai
  • #cybersecurity
  • #llm

Related posts