deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Artificial Analysis Intelligence Index v4.2 adds harder agentic tasks, drops saturated GPQA Diamond

Artificial Analysis has updated its Intelligence Index to v4.2, adding agentic knowledge work and long-document benchmarks, doubling private held-out test weighting to 40%, and retiring the saturated GPQA Diamond.

Artificial Analysis Intelligence Index v4.2 adds harder agentic tasks, drops saturated GPQA Diamond

Artificial Analysis has shipped version 4.2 of its Intelligence Index, an interim release that pulls parts of the planned v5 update forward in response to rapid movement at the AI frontier. The company says it had deliberately held back changes to keep the index stable through recent major model launches, but decided an immediate update was warranted — it has been eight months since v4 launched in January. The revision, announced in a post that reached the Hacker News front page, adds harder and more realistic tasks and doubles the share of scoring drawn from private test sets.

New agentic and document benchmarks

The headline addition is AA-Briefcase, an in-house evaluation of agentic knowledge work built on a private, held-out test set. According to Artificial Analysis, models are tested on multi-week projects designed by industry experts, each involving many linked tasks and thousands of input source files. Scoring combines rubric and pairwise grading to measure verifiable task success, analytical quality and presentation quality, aiming to capture how well a model handles sustained knowledge work rather than isolated questions.

The second new benchmark, GDP.pdf, was created by Surge AI. It tests single-turn professional document reasoning across 100 PDFs in ten domains, asking models to synthesize evidence spread over 4,592 pages of text, tables, charts, footnotes and exclusions. Responses are graded against 1,275 expert-authored atomic criteria, and the headline All-pass Rate credits a task only when every criterion is satisfied.

GPQA Diamond, meanwhile, has been dropped. Artificial Analysis calls it an exceptional scientific reasoning evaluation that has now been saturated — effectively maxed out by frontier models, so it no longer separates them.

Less room for gaming

Forty percent of the index's weighting now comes from private, held-out test sets, double the share in v4.1. The held-out data includes AA-Briefcase, AA-Omniscience and solutions for CritPt. The stated goal is to reduce labs' ability to game evaluations, since public benchmarks can leak into training data whether intentionally or not. Artificial Analysis says the held-out percentage will rise again in v5.

Grading infrastructure updates

The update also reworks how responses are scored. AA-LCR v1.1 gains a grading system prompt, and errors and ambiguities in its answer keys have been corrected. For GDPval-AA v2 and AA-Briefcase, sampling was improved and the Elo scale re-anchored so ratings stay stable as new models are added. SciCode's grading sandboxes were made more robust so that slow but correct code no longer counts as a failure.

What the new leaderboard shows

Anthropic's Claude Fable 5.1 now leads the index, followed by OpenAI's GPT-6 Astra, which posts a 4-point gain over GPT-5.6 Sol. Meta is the third-ranked lab, ahead of SpaceXAI, Moonshot/Kimi, Z.AI and Google.

The two new benchmarks split the field. On AA-Briefcase, Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3, with GPT-6 Astra sitting roughly 85 Elo points above GPT-5.6 Sol. On GDP.pdf, OpenAI leads: GPT-6 Astra scores a 33.2% All-pass Rate and GPT-5.6 Sol 28.2%, with Claude Fable 5.1 at 26.2%.

On cost, four labs — Anthropic, OpenAI, Meta and Z.AI — share the updated Cost per Task Pareto frontier. GPT-6 Astra is more token-efficient than nearly every other model near the intelligence frontier, according to Artificial Analysis, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at the extremes of the efficiency curve (models scoring below 25 on the index are excluded).

Why it matters

Public benchmarks are losing their ability to discriminate as frontier models saturate them, and benchmark gaming has become a live concern for anyone using leaderboards to pick a model. Artificial Analysis's move toward private, agentic, multi-week tasks is a bet that the next useful signal will come from held-out evaluations that resemble real work. For buyers, the methodological shift matters as much as the rankings: because the test mix and weightings have changed, score movements partly reflect the new methodology, and the divergence between AA-Briefcase and GDP.pdf results shows that the answer to which model is best increasingly depends on the task. More incremental releases, and the full v5, are planned next.

  • #llm-benchmarks
  • #ai-evaluation
  • #artificial-analysis
  • #frontier-models