deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

DoGBench: first benchmark for user-facing docs finds no agent clears 55 of 100

DoGBench scores AI-written documentation patches against maintainer-validated rubrics. The best system reaches 54.8/100, with missing prerequisites and subtle hallucinations among common failures.

DoGBench: first benchmark for user-facing docs finds no agent clears 55 of 100

A benchmark for documentation work

A new benchmark called DoGBench, which reached the Hacker News front page, claims to be the first to evaluate AI agents on user-facing software documentation — the guides, references and tutorials that end users actually read. Rather than testing generic writing ability, it measures whether an agent can recognise when documentation needs updating and produce a patch that would survive expert review.

According to the research published at dogbench.ai, the benchmark contains 292 tasks drawn from real open-source projects including Helm, PostHog, Mautic and Doc Detective. Of these, 205 require a documentation change and 87 require leaving the documentation alone. The headline evaluation uses a stratified 117-item held-out split: 82 tasks needing an update and 35 needing none.

Each agent receives the repository as it existed before a change, plus a triggering signal such as a code pull request or a reported documentation gap. The merged documentation is withheld. Patches are scored against task-specific rubrics validated with project maintainers, covering accuracy, completeness, reader guidance, placement, style and repository conventions. Across the 205 documentation-needed items, the rubrics contain 3,273 criteria, of which 798 are P0 — failures that cap a patch's score at 60 out of 100 regardless of other strengths.

Low scores across the board

The results are sobering. Among the seven model-and-harness lanes reported in the paper, the best combined score was 47.3 out of 100, achieved by Qwen3.8 Max running with OpenCode. The public leaderboard adds three cloud-agent systems evaluated on the same split, and even the top entry there, a cloud agent, reaches only 54.8. The DoGBench team notes the scores measure progress toward an expert standard, not performance relative to a human expert.

Reliability is a bigger problem than raw quality. The highest P0-clean delivery rate among the paper's seven lanes was 39.0%: GPT-5.6 Sol with Codex produced a patch with no critical failures on just 32 of the 82 held-out tasks requiring an update.

Where agents fail

The research identifies two recurring failure modes. The first is judgement: deciding whether a change affects users at all. Agents err in both directions — some write documentation for internal refactors or performance work that needs no new user guidance, while others skip necessary updates because they cannot find an existing page about the feature, even when that missing coverage is precisely the problem they should solve.

The second is completeness. A patch can read clearly and contain correct facts yet omit a prerequisite, a decision point or an essential step. In a broader audit of 1,267 submissions, 45.5% had a task-completion gap such as a missing prerequisite, procedural step, verification step or recovery path; 36.6% contained technical inaccuracies; and 32.5% omitted part of the central concept or reference information. These categories overlap. Outright fabrication — invented classes, flags or endpoints — appeared in 6.1% of submissions, and the team flags subtler distortions from frontier agents: extending the scope of real functionality, misstating defaults, or presenting behaviour that holds only under specific conditions as universally true.

Trajectory analysis adds another finding: agents frequently stopped investigating after finding the first plausible page to edit, missing decisive evidence elsewhere in the repository.

Scoring design choices

DoGBench deliberately separates two skills. Delivered patch quality averages quality across the 82 update tasks, with missed or empty patches scoring zero. Correct abstention measures how often an agent leaves documentation unchanged across the 35 no-change tasks. The combined score is the harmonic mean of the two, 2DN/(D+N), so strength on one dimension cannot compensate for weakness on the other. For reproducibility, the team is publishing 735 complete agent trajectories from the seven non-cloud agents. The work is credited to Frances Liu and colleagues as a 2026 manuscript.

Why it matters

Documentation is a plausible near-term use case for coding agents, yet this benchmark suggests it is far from solved. Even the best systems leave readers unable to complete tasks roughly half the time by the benchmark's own rubrics, and the P0-clean rates imply that expert review remains mandatory. A documentation owner's context — who the reader is, what they are trying to accomplish, how a change affects their workflow — is exactly what agents lack. For teams evaluating agent tooling, DoGBench offers a concrete yardstick and a reminder that fluent prose is not the same as useful guidance. The benchmark is accepting external submissions, which are reviewed manually before results appear on the leaderboard.

  • #ai-agents
  • #benchmark
  • #documentation
  • #llms
  • #open-source

Related posts