deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Microsoft and Hugging Face's ThinkingBox grades AI agents on terminal state, not words

Microsoft and Hugging Face have reportedly published ThinkingBox, a benchmark that grades AI agents on the backend state and side effects they leave behind rather than the text they generate.

Microsoft and Hugging Face's ThinkingBox grades AI agents on terminal state, not words

A benchmark that scores the system, not the transcript

Microsoft and Hugging Face have published ThinkingBox, a benchmark that grades AI agents on the backend state and side effects they actually produce rather than the sentences they generate, according to a write-up on dev.to by Igor Eduardo. As described in the post, the benchmark spans 507 stateful workflows and runs each task 20 times, testing whether an agent can reliably reach the correct end state rather than describe success once. Eduardo is careful to note he has not run the benchmark himself and that those figures belong to its authors.

The failure it is built around

The benchmark's opening example, as relayed in the post, is a support agent that makes nine well-formed tool calls and marks a ticket as resolved, and is wrong: a carrier exception was still open, so the required end state was "on hold." A grader that reads the tool calls sees nine clean invocations. The database disagrees.

Eduardo's broader argument is that most agent dashboards answer the question "did it say it succeeded?" Status strings, traces and the final summary are all produced by the same run that took the action, which makes them useful for debugging but useless as a grade. When the evaluation and the agent share a witness, a confidently wrong action scores identically to a correct one, and the failure resurfaces later as a customer complaint, a ledger discrepancy or an audit finding. A companion dev.to piece, titled "The Witness Was the Suspect," makes the same point from the incident side: when an agent writes its own success log, the record of an incident is authored by its cause.

Separating the actor from the grader

Eduardo sketches five questions an outcome-based evaluation should answer. First, a required end state written down before the run, specifying what the record should look like when the task is truly finished. Second, an independent read path, so the final state is read through a query the agent cannot influence. Third, a side-effect ledger covering collateral changes such as tickets, refunds, emails and database rows, so a correct primary record with wrong surrounding writes does not pass silently. Fourth, a separately counted failure mode for runs that claimed success while the state check failed. Fifth, repeat consistency, checking whether the same task reaches the same end state across multiple runs.

None of this requires a private harness, he argues. It requires someone to specify the required end state before the run and a grader that reads the system of record after it.

Why twenty runs matter

The repeat-consistency framing is the part of ThinkingBox Eduardo singles out. A single green run shows an agent can reach the right state; production needs to know whether it will. If the same ticket ends resolved on some runs and on hold on others, the result is not a capability with a small error rate but unpredictable behavior hidden behind a convincing log. He argues consistency should be reported alongside the pass rate rather than relegated to an appendix.

The same split in retrieval

Eduardo, who builds retrieve-first systems, draws a parallel to RAG evaluation. A generated answer can cite sources cleanly and still be wrong about them, just as an agent can log cleanly and leave the wrong record. The remedy is the same in both cases: grade against something the generator did not produce. For retrieval that means a gold source set; for agents it means the terminal state.

Why it matters

The details here come from a third-party post rather than the primary announcement, but the reported design points at a real shift in how agents could be measured. Moving the grade from self-reported success to verified system state closes the gap between what an agent claims and what it actually did, and naming the false-success failure mode turns a quiet production risk into a measurable metric. The practical takeaway for teams is compact: define the required end state first, score from a system of record the agent cannot touch, count claimed-success-while-wrong as its own category, and run the task more than once before believing the result.

  • #ai-agents
  • #benchmarks
  • #evaluation
  • #llms
  • #hugging-face

Related posts