deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Epoch AI's InnovationEval: frontier models struggle to invent novel ML techniques

Epoch AI's new InnovationEval benchmark asks whether AI can independently rediscover a recent post-training innovation. Frontier models made little progress despite thousands of dollars of GPU time.

Epoch AI's InnovationEval: frontier models struggle to invent novel ML techniques

What InnovationEval measures

Epoch AI has released early results from InnovationEval, a benchmark that tests whether an AI system can independently discover a novel machine learning technique matching the performance of a recent human-developed innovation it has never seen. The answer, so far, is not much: according to Epoch AI, recent frontier models make little progress on the task even when allowed to spend thousands of dollars of GPU compute on experiments. The publication circulated on Hacker News's front page.

The question behind the eval is a concrete version of a classic ideation thought experiment — whether an AI armed with humanity's knowledge up to 1905 could rediscover special relativity. Epoch AI poses a more modest variant: given what AI researchers knew in early 2026, can a model find its own algorithmic improvement on par with what humans have published since? The work sits alongside other end-to-end research benchmarks such as Crux Evals, ResearchGym and RSI-Bench.

The task and how it is scored

The testbed is on-policy self-distillation (SDPO), a post-training technique that Epoch AI says has been adopted and cited by recent models. The agent under evaluation must invent a new post-training method that beats a strong GRPO baseline, demonstrating it by fine-tuning a Qwen3-8B model on short-answer questions and coding. It is asked to produce evidence that ideally matches or surpasses reference values set by the original, unnamed method, in the form of a result convincing enough to move the field forward.

Scores are averaged across the two task areas, each built from several sub-metrics taken from the original paper's experiments. Matching or beating the paper's numbers in an area yields 100%, while performing at the GRPO baseline yields 0%.

Forcing genuine novelty

Epoch AI argues that many existing AI R&D evaluations examine well-specified tasks that require no innovation, or that can be solved by recombining known techniques. InnovationEval is designed so that substantially improving the metrics within scope requires a method the model has not seen during training — though the discovered technique need not resemble SDPO itself.

Enforcing that scope is delicate. Left unconstrained, an agent could inflate the metrics through shortcuts such as generating synthetic fine-tuning data, which would raise benchmark numbers without constituting a new post-training technique. Epoch AI therefore restricts the agent to algorithmic changes that affect the loss and its updates, or the rollouts and model-driven revisions applied to a fixed batch of training data. The organization acknowledges this invites a dynamic where the agent keeps hunting for loopholes and reviewers disallow solutions after the fact, but says even imperfect scope limits reduce the burden of reviewing submissions.

Grading could not be fully automated in this iteration. Epoch AI experimented with an automated judge, Opus 5, to detect scope violations, but ultimately relied on human-in-the-loop review of each submission's results and workings. Each attempt also demands substantial compute, which limited the evaluation to a small number of runs.

Contamination is already a problem

Anchoring the eval on a real published method guarantees the task is feasible and clarifies the required budget, but it also means newer models can simply memorize it. Epoch AI says this happened over the course of the project: Claude Fable 5 and GPT-5.6 Sol showed no sign of knowing the paper's details when asked to guess without search, but their successors, Claude Fable 5.1 and GPT-6 Astra, were already aware of the task. Future rounds will test for memorization, flag affected results and swap in fresh tasks as needed.

Why it matters

Building an automated AI researcher is a stated goal of leading AI developers, and AI has already shown competence at pieces of research work: software engineering tasks, dataset creation and open-ended optimization of defined metrics. InnovationEval targets the missing middle — the full loop of generating ideas, implementing them, running experiments, analyzing results and iterating to something genuinely new. On that measure, current frontier models remain far from automating AI R&D itself. That is a useful, trackable signal for anyone estimating when AI might meaningfully accelerate its own development, and a data point on how much of research work stays human for now. Epoch AI plans to expand and repeat the methodology to chart progress over time.

  • #ai
  • #benchmarks
  • #machine-learning
  • #evaluation
  • #llms

Related posts