· via dev.to (home feed)
Google DeepMind launches Gemini 4 Argon with million-token output and defenders-first access
Google DeepMind's Gemini 4 Argon claims first place on most launch benchmarks, can emit up to a million tokens per call, and undercuts Claude Opus 5.5 on price — but initial access is limited to vetted cyber defenders.

Google DeepMind announced Gemini 4 Argon on September 30, 2026, its first new frontier model since Gemini 3.1 Pro last November and its first release outside the Flash line in roughly seven months. According to a dev.to analysis of the launch, Argon takes first place on 13 of 18 benchmarks measured against GPT-6 Astra and Claude Opus 5.5, a tally the post credits to VentureBeat. It raises the per-call output limit from around 64,000 tokens to about one million, and its introductory API price is half of Claude Opus 5.5's. Almost nobody can use it yet: the first wave of access is going to vetted cybersecurity defenders.
What the benchmark sheet says
Google's launch numbers, as relayed by the dev.to write-up, put Argon first on long-horizon software engineering (DeepSWE v1.1 at 77.9%, against 74.2% for Opus 5.5 and 74.1% for Astra), end-to-end business workflows (AutomationBench at 51.3% versus 42.5% and 41.4%) and legal research and drafting (Harvey Legal Agent at 19.6% versus 5.4% and 3.8%). It also leads FrontierSWE V2 at 55%, LAB-Bench 2 at 88.8%, LVBench long-video understanding at 91.7% and Vals Finance Agent at 65.4%.
On security, the model scores 68% on CWE-bench v1, tied with GPT-6 Astra, and attackers succeeded in only 0.7% of attempts on the Gray Swan prompt-injection suite, the best reported figure. Artificial Analysis puts its Intelligence Index at 53, level with GPT-6 Astra, but records Terminal-Bench 4 at 57%, behind Sonnet 5.5 at 64%, Opus 5.5 at 60% and Astra at 59%, so the lead in agentic coding is not universal.
The million-token output ceiling
An input context of a million tokens is no longer the differentiator; the output side is the change. Previous Gemini models capped output near 64K tokens. Argon can emit roughly a million in a single call — the dev.to post compares that to about 3,000 pages of text, or a mid-sized codebase rewritten without stopping. The practical argument is about agents: a model that can hold an entire plan and its full deliverable in one generation spends less time chunking work, resuming and re-establishing context, which is where long-running tasks tend to drift. Google disclosed almost nothing about architecture or parameter count. Argon is described as a reasoning model with extended thinking, aimed at software engineering, legal and finance knowledge work, and cybersecurity.
Half the price of Opus, at first
The introductory rate is $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 (a 95% discount), before reverting to $4/$20. That opening price is half of Claude Opus 5.5's $4/$20 list and about a fifth of GPT-6 Astra's pricing. The source frames this as a trend rather than a one-off: Opus 5.5 itself cut 20% off input and output and 60% off cached reads relative to Opus 5, so frontier capacity is being repriced for volume.
Defenders get it first
Access is staged. Vetted cyber defenders in Google's Fairwind Program receive Argon before anyone else, and the US government gets pre-release access under a voluntary process. Thousands of Google employees already use it internally; one cited example had the model beat a published baseline by 40% while optimising quantum computing subroutines. Paid API customers and Google AI Ultra subscribers are said to be next, with no date attached.
Caveats worth holding
Bloomberg reported, citing anonymous sources, that some Google employees considered the model's internal performance inadequate; Google disputed the account. Google's methodology note also concedes that some Argon scores were computed internally while rival scores came from provider publications or public leaderboards, so the margins should be read as directional. Even a category-leading 19.6% on Harvey remains modest in absolute terms.
Why it matters
The launch choreography may matter more than the leaderboard. Handing a capable vulnerability-finding model to defenders before general availability treats dual-use risk as a rollout problem rather than a policy footnote, and other labs are likely to copy the pattern. The pricing signals that frontier capacity is increasingly sold as a volume business. The benchmark contest has also moved to professional workflows — legal, finance, automation — where Argon's margins are widest. Until access widens and independent evaluations run, though, every number here rests on Google's own reporting and a handful of secondary accounts.
- #gemini
- #google-deepmind
- #llms
- #benchmarks
- #api-pricing