deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

livenerf: open append-only benchmark tests whether Claude Opus 5.5 degrades post-launch

A new open-source benchmark runs a calibrated 78-question panel daily against Claude Opus 5.5's launch-week baseline, aiming to turn recurring post-release 'nerf' claims into measurable data.

livenerf: open append-only benchmark tests whether Claude Opus 5.5 degrades post-launch

A launch-day clock for the nerf debate

An open-source project called livenerf has started a 30-day measurement campaign against Anthropic's Claude Opus 5.5, which the repo says was released on 22 September 2026. Its single goal: determine whether a model's capability drifts after it ships. The project, hosted on GitHub and surfaced on Hacker News's front page, was motivated by months of community claims that Anthropic quietly degrades shipped models through quantization, a smaller model behind the same name, reduced inference effort, or routing changes. The README is candid about the rival explanation — nothing happened and users are reading patterns into noise — and notes that, without a clean baseline recorded at launch, the argument has consisted of unverifiable claims from both sides.

How the harness stays honest

The first run was on 24 September 2026, roughly two and a half days after launch, and the same panel executes once a day for 30 days. Days 1–10 form the baseline, followed by two 10-day comparison windows, so the earliest verdict could arrive around 24 October; the first results row appears after day 20.

Because the model itself cannot be made deterministic — sampling parameters are unavailable and thinking cannot be switched off — livenerf pins everything around it instead: frozen prompts, a pinned CLI version, exact graders, and raw logs retained permanently. Drift is then assessed statistically across thousands of samples. The harness is built on Inspect, the UK AI Security Institute's open-source evaluation framework, and the statistics follow Anthropic's own "Adding Error Bars to Evals" guidance, so the methodology leans on established practice rather than bespoke choices.

The panel was calibrated rather than assembled by hand. According to the repo, 2,336 questions drawn from GPQA Diamond, MMLU-Pro, competition math and AIME 2025–26 were each sampled four times; Opus 5.5 answered about 93% correctly on the first try, and 97% of questions were either always right or always wrong. The 78 questions it answers inconsistently became the panel. The project also quantified its own selection bias — on fresh samples the pass rate rose from 54.7% to 62.0% — and used the fresh figures for its power calculations.

Unusually, v0 runs on a Claude Max subscription through headless Claude Code rather than an API key, with budgets expressed as a share of the plan's weekly usage meter. One run per day of the full panel can detect an accuracy change of roughly 7.5 points per 10-day window while consuming about 3.6% of the weekly plan.

What it can and cannot detect

As of 29 September, six of 30 days had been collected with none missed, all running the full 90 samples on the same harness hash and pinned CLI 2.1.280, with a single budget-guard override logged for day five.

Validation passed its pre-registered criterion, and it showed where degradation becomes visible first: token counts. Running at low effort cut output tokens by 62% while costing 8.3 ± 4.5 accuracy points; medium effort cut tokens 26% for 4.2 ± 3.9 points. Median output tokens per sample are tracked as a secondary signal precisely because quietly reduced thinking would likely show up there before accuracy moves at all.

The stated limits matter as much as the promises. A validation that swapped in Opus 5 for Opus 5.5 produced a difference of −3.8 ± 6.3 points and −23% tokens — not distinguishable at 99% confidence. The instrument therefore cannot rule out a same-family model swap of that magnitude at validation sample sizes; a full 10-day window contains roughly 2.5 times the samples, but the repo says sufficiency has not been demonstrated.

An audit of the panel questions found 8 wrong answer keys and 30 ambiguous items out of 80 reviewed. Nothing was dropped; a pre-registered sensitivity analysis reruns the results without them. Where Anthropic's serving path sometimes answered with Opus 5 or refused certain biology and math questions, those samples were rejected and counted, and the affected questions excluded.

Why it matters

Claims that shipped models get silently worse are now a recurring feature of frontier-model discourse, and they are effectively unfalsifiable without a launch-day baseline. livenerf's contribution is procedural as much as technical: a pre-registration committed to a public git timestamp, an append-only results log, clustered standard errors, and a commitment to report gains with the same prominence as regressions. Even a null result would be informative, and the honest disclosure of detection limits — including that a modest same-family swap might slip through — sets a standard other evaluators rarely meet. Because the design is clonable and budget-aware, it also offers developers a reusable template for independent monitoring of any model they depend on, not just this one.

  • #anthropic
  • #claude
  • #llm-evals
  • #benchmarks
  • #open-source

Related posts