deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Fable 5 tops Prime Intellect's 18-model nanoGPT speedrun, closing 81.7% of human record gap

Prime Intellect benchmarked 18 frontier models with 153 autonomous runs on the nanoGPT optimizer speedrun. The best model closed 81.7% of the gap to the human record, and 41 agent traces are now public.

Fable 5 tops Prime Intellect's 18-model nanoGPT speedrun, closing 81.7% of human record gap

Prime Intellect has published a large autonomous benchmark that put 18 frontier models to work on the nanoGPT optimizer speedrun, a community challenge built around training a small GPT model as efficiently as possible. Across 153 autonomous runs, the strongest performer — a model listed as Fable 5 — closed 81.7% of the distance between the benchmark's baseline and the human record. No model beat the record. The research page surfaced on the Hacker News front page.

How the results are measured

Runs are scored on the speedrun's metric, where lower numbers are better: the page puts the human record at 2,600 and the unmodified baseline at 3,290. For every model, Prime Intellect reports its best validated score and how much of the gap to the human record that run closed. Validation matters at the bottom of the table — GLM 5.3 is listed with no record at all.

What the leaderboard shows

Fable 5 stands alone at the top with a score of 2,726, an 81.7% gap closure, produced with the claude-code harness at high effort over 8.7 days of agent time. It is the only entry past the halfway point. Opus 5 is second at 2,920 (53.6%), followed by Kimi K3 at 2,930 (52.2%) running under the prime-agent harness. Kimi K3 appears twice on the board; on its own kimi-code harness it reached only 2,974 (45.8%) over a longer 5.1-day run — evidence that scaffolding choices move results nearly as much as model choice.

The mid-field clusters between roughly 11% and 40%: Opus 4.8 at 39.4%; the GPT-5.6 family spread from Sol (35.9%) and Sol Pro (33.6%) down to Terra (11.0%); Sonnet 5 at 26.8%; Grok 4.5 and Qwen3.8 Max tied at 24.6%; GLM 5.2 at 20.3%; and DeepSeek V4 Pro at 12.3%. The tail — Grok 4.6, Muse Spark 1.2 and 1.1, GPT-5.5 and Kimi K2.7 — all closed under 11%, with Kimi K2.7 last among validated runs at 7.2%.

Time and harness budgets

Agent time varied more than tenfold, from 0.6 days for Grok 4.6 and Muse Spark 1.2 to 8.7 days for Fable 5's winning run. Several mid-table models — Sonnet 5, GPT-5.6 Luna, Qwen3.8 Max, GLM 5.2 — finished in roughly two days. The runs used a range of harnesses (claude-code, codex, prime-agent, kimi-code, grok-cli, qwen-code, pi and muse-code) at effort tiers from high up to a level labelled xhigh.

Because raw run length and compute differ across entries, the page also offers an equal-budget view. It hands each model's best final run the same resource budget — adjustable from 6 hours to 9 days, with a 24-hour default — and reports the best validated record reached inside that window; runs that stopped before the chosen budget are greyed out.

Open traces

Alongside the leaderboard, Prime Intellect released 41 curated, full-length agent trajectories covering tool calls, subagent activity and scratchpads. That means the difference between a 24.6% run and an 81.7% run can be inspected step by step rather than taken on faith.

Why it matters

Benchmarks for agentic coding abound, but controlled, cross-model data on multi-day autonomous research work is scarce. This one offers three things: a like-for-like capability ranking on an open-ended optimization task, a documented harness effect (the same Kimi K3 model lands several points apart depending on its tooling), and a floor-to-ceiling view of how far frontier agents still sit from expert humans — even the best run left 18.3% of the gap unclosed. The released traces may be the most durable output: they turn a leaderboard into a dataset, giving researchers raw material for studying long-horizon agent behaviour and giving future training runs concrete examples to learn from.

  • #ai
  • #benchmarks
  • #llm-agents
  • #nanogpt
  • #machine-learning

Related posts