· via Hacker News – Front Page (native)
Artificial Analysis sizes up Gemini 4 Argon (High) on intelligence, cost and token use
Artificial Analysis has published its Intelligence Index breakdown of Google's Gemini 4 Argon (High), spanning agentic benchmarks, hallucination scoring, token use and per-task pricing. The page surfaced on Hacker News.
Artificial Analysis has published an updated intelligence, performance and price assessment of Google's Gemini 4 Argon (High), giving developers a single reference point for weighing the model's capabilities against what it costs to run. The model page, dated 30 September 2026, was picked up on the Hacker News front page.
How the Intelligence Index is scored
The centrepiece is the Artificial Analysis Intelligence Index v4.3.2, which rolls ten evaluations into one number: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. According to Artificial Analysis, the suite leans toward agentic work rather than static question answering. It spans agentic knowledge work, agentic real-world work tasks, agentic coding and terminal use, professional document reasoning, scientific research workflows run in a terminal, and medical long-context reasoning under the MLCR-AA label. The index also charts open-weight models separately from proprietary ones, and flags licences that restrict commercial use. Alongside the headline score, capability indexes slice performance by specific skills and industries, on the grounds that general intelligence does not translate evenly to every use case.
Agentic benchmarks and Elo scoring
AA-Briefcase v1.1, an agentic knowledge-work benchmark developed by Artificial Analysis itself, reports an Elo figure that blends three signals: a rubric pass rate, analytical quality and presentation quality. Rubric results are converted into Elo through simulated head-to-head comparisons, and both Elo values and their confidence-interval bounds are clamped at zero. The remaining evaluations round out the picture, from terminal-driven coding to long-context reasoning over medical material.
Hallucination and calibration
AA-Omniscience measures knowledge reliability and hallucination. Correct answers push the score up, hallucinated answers pull it down, and refusing to answer carries no penalty. Scores range from -100 to 100, where zero means as many incorrect answers as correct ones and negative scores indicate more wrong than right. That asymmetry is deliberate: it separates models that guess confidently from models that know when to abstain, a distinction that matters far more in production than in a quiz-style test.
The economics of running it
The pricing side of the analysis is expressed as a weighted average cost, in dollars, per Intelligence Index task. For each of the ten evaluations, Artificial Analysis takes the model's input, cache-hit, cache-write, reasoning and answer token prices, divides by task count, and weights the result by that benchmark's share of the index. The page also reports the total cost of running every evaluation in the suite, the weighted average number of output tokens consumed per task, per-million-token prices split across cached prompts, regular input and output, and the context window, meaning the maximum combined input and output tokens, with output typically capped well below that limit. Cached prompts are typically offered at a significant discount relative to regular input, but cache-write and cache-storage charges are billed separately and vary by provider. Larger context windows matter mainly for retrieval-augmented workflows that reason over large document sets.
The specific index scores and dollar figures sit on the model's analysis page alongside this methodology, which is what makes the comparison repeatable across models.
Why it matters
Choosing a frontier model is now as much a unit-economics decision as a capability decision. Expressing price as cost per completed benchmark task, with reasoning tokens included in the arithmetic, lets teams compare a cheap-but-verbose model against an expensive-but-efficient one on equal footing. The benchmark mix also signals where evaluation is heading: toward multi-step agentic work in terminals, documents and research workflows, rather than single-turn knowledge questions. And the hallucination penalty, paired with a free pass on refusals, rewards calibration, which is precisely what production deployments need. For high-volume applications, the split between cache-hit, cache-write and storage pricing can dominate the bill more than headline input and output rates, so Google's caching terms deserve as much scrutiny as Gemini 4 Argon's intelligence score.
- #gemini
- #llm-benchmarks
- #artificial-analysis
- #ai