deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Microsoft's ThinkingBox finds AI agents pass all 20 attempts on fewer than half of tasks

A Microsoft/Hugging Face benchmark that grades database state instead of chat transcripts found top agents complete all 20 runs of a task on barely half of business workflows, and two-thirds of failures raise no error.

Microsoft's ThinkingBox finds AI agents pass all 20 attempts on fewer than half of tasks

A benchmark that grades state, not transcripts

Microsoft researchers and Hugging Face have published the October results of ThinkingBox, a benchmark built around a single design choice: it never reads the agent's chat transcript. Instead, it snapshots a backend database before the agent starts, lets the agent work through its tools, and then uses deterministic judges to compare the final database against a required end state and to flag forbidden side effects on records that should have stayed untouched. The accompanying paper, according to a dev.to write-up, is titled "One Success Isn't Reliability," which is the whole argument in four words.

The benchmark covers 507 business workflows across retail, travel and hospitality, auto insurance, neobank IT and consulting IT/HR — tasks such as modifying orders, filing insurance claims, rebooking travel and resolving tickets. Tools are exposed as MCP servers, so any agent that speaks the Model Context Protocol can plug in. Each run gets a freshly reset, isolated backend so one attempt cannot contaminate the next, and every task is repeated 20 times.

Results are reported three ways: pass@1 for first-attempt quality, pass@20 for the capability ceiling (solved at least once in twenty tries), and observed 20-of-20 for dependability (solved every single time).

The numbers

Across 121,680 trials, the benchmark recorded 79,853 failures. The leading model, Claude Opus 5.5, passed all 20 attempts on just 241 of 507 tasks — 47.53% — while posting 66.50% at pass@1. Kimi-K3 fell harder, from 57.37% pass@1 to 17.60% at 20-of-20. The Hugging Face write-up frames the Kimi result starkly: it solves 93.89% of tasks at least once but is dependable on only 13.41%. The dev.to author cautions that exact decimals differ slightly between the paper abstract and the HF post, though the direction of every finding is consistent.

Domain matters as well. Retail workflows averaged 59.52% pass@1, while auto insurance averaged 33.83%, which the write-up attributes to insurance carrying more policy rules and more fields that must be exactly right — precisely the kind of task a business would want to automate.

Where the failures come from

Two findings stand out for anyone running agents in production.

First, 67.24% of failed attempts ended with no tool error at all. No crash, no loop, no rate limit — just a confident summary produced after valid state-changing actions. Monitoring that watches for errors and timeouts would have marked these runs as successful, and the mistake would surface later, from a customer, an auditor or a reconciliation job.

Second, root causes cluster heavily around tool handling: 79.9% of failures traced to the agent picking the wrong tool, misreading a schema or fumbling parameters. Wrong updates accounted for 10.3%, incomplete resolution for 7.0%, and no action at all for 2.9%. By failure type, roughly 77.6% of failed trials involved wrong field values, 43% created unintended side effects and 25% missed required changes — categories that overlap, since one trial can fail several ways.

The write-up also estimates cost per dependable task at about $6.80 for GPT-5.4, $7.45 for GPT-6 Astra and $7.80 for Claude Opus 5.5, arguing that a cheaper model needing more retries or human cleanup can cost more per finished task than its per-token price suggests.

Caveats

The dev.to post is explicit about the limits. The 507 workflows are well-designed but synthetic stand-ins for real systems, all results come from a single agent harness, and the deterministic judges are only as good as the end states someone wrote by hand. Microsoft has released the framework under an MIT licence on GitHub with the dataset on Hugging Face, and the authors encourage others to reuse the method rather than just read the leaderboard.

Why it matters

The benchmark quantifies the gap between what an agent can do on a good day and what it does every time, and that gap is the real adoption question for agentic AI. Demos are single runs; production workflows such as refunds, bookings and claim payouts must succeed repeatedly, and pass@1 systematically overstates what can be promised for them. Because two-thirds of failures are silent, error-rate dashboards cannot catch them — only checks on final state can, which argues for putting reconciliation-style verification into production monitoring for any agent with write access. And with tool handling behind 79.9% of root causes, sharpening tool schemas, descriptions and parameter naming may be a higher-leverage reliability fix than swapping models.

  • #ai-agents
  • #benchmarks
  • #mcp
  • #reliability
  • #microsoft
  • #hugging-face

Related posts