· via dev.to (home feed)
Microsoft's Decision-1 probability model works as a cheap verifier for AI tool calls
A dev.to post describes using Microsoft's Decision-1, a probability-output model on OpenRouter, to verify AI tool calls for thousandths of a cent each, as long as it is asked facts rather than judgments.

A model that only outputs probabilities
Microsoft recently listed a small, specialized model on OpenRouter under the name Microsoft-Decision-1. Unlike a general-purpose chat model, it accepts a described situation plus a question with a fixed set of named answers, and returns a probability for each option. According to a dev.to post by Tom Jones of Spanda Works, that is all it does.
The Spanda Works team runs a gateway that sends most requests to inexpensive models and inspects the results before they go out. A component that reduces any question to a yes or no with a confidence number looked like the verification layer they had been trying to build, so they spent a day constructing a checker around it.
Steady on facts, unsteady on judgment
The first experiments, per the post, were discouraging. Asked the open questions you would pose to a human reviewer, such as whether a note is relevant to an agent's next step, the model returned inconsistent probabilities. Identical relevance queries varied by as much as 0.17 between calls.
Where it held firm was on verifiable facts. A question like whether a user's message contains a specific city name came back the same, and correct, every time. That observation set the design rule for the day: deterministic code establishes the facts, and Decision-1 gets asked only clear-cut questions about facts already laid in front of it. Questions requiring real judgment return UNSURE and escalate to a stronger model.
The checker pipeline
The resulting system is deliberately plain. Code validates each tool call, Decision-1 answers a handful of single-fact questions, and if something fixable is wrong, a cheap repair model gets one attempt. The repaired call then passes through the same checks. Every run ends in APPROVE, REJECT or UNSURE, with a record of every check that executed.
One late rule mattered most: a repair is kept only if the checker approves it afterward. Otherwise the original call goes out untouched, which is why the final version broke nothing that was already correct.
Measured results across four versions
The team defined three failure modes and locked pass and fail criteria written in advance by a separate model that had not seen their results. Each version ran once on 300 tool calls it had never encountered, scored by fresh AI labellers working from a written rubric, blind to whether a call was original or repaired.
Wrong approvals fell from 9 of 48 in version 2.1 to 0 of 72 in version 2.4. Wrong calls that were repaired and approved moved from 25 of 100 to 20 of 96. Calls broken by a repair went from 2 of 200 to 0 of 201. Most of the improvement came from studying earlier failures and adding checks for shapes code can see, such as a list serialized as text or a user asking for fewer than five items and receiving a limit of exactly five.
The post's own reviewer model pushed back on the zero, noting that each version ran on a different draw of calls with different labellers, which adds noise, and that the fix rate of 20 out of 96 clears their 20 percent floor only narrowly. The author counters that the raw counts stand on their own: across the four versions the checker approved 164, 150, 120 and 110 calls, and was wrong on 9, 4, 2 and 0 of them. The tradeoff is that version 2.4 says yes far less often.
Cost per task, not per token
Decision-1 is inexpensive enough to be nearly invisible: the full 300-call test cost under a cent, and on 1,417 additional recorded calls it ran at roughly two thousandths of a cent per call with a median latency under a second.
The more interesting shift, the author argues, is what a checker changes about pricing. Without one, you pay per token and hope the answer is right. With one, the unit becomes a verified answer, and its price includes every call the checker refuses to approve and must send somewhere more expensive. On their recorded runs, a cheap model plus the checker plus escalation cost a fraction of running the same tasks directly on a premium model, with the size of that fraction governed almost entirely by how often the checker can say yes.
The remaining gap
About four in ten calls still come back UNSURE, mostly cases where correctness depends on what a value means rather than whether it is present. Code cannot inspect meaning, and the decision model asked to judge it becomes inconsistent. The question the team is carrying forward is how to convert a judgment about meaning into a fact that plain code can verify, without teaching the checker to approve things it should not.
Why it matters
Verification is the main obstacle to trusting cheap models with production work. This writeup suggests a workable pattern: a specialized probability model as a fast, near-free fact-checker, with deterministic code supplying stability and a stronger model absorbing ambiguity. Zero wrong approvals at thousandths of a cent per check shows that oversight costs can drop low enough to make inexpensive models viable for real tasks, provided you accept that a large share of calls still needs escalation, and that precision is bought by approving less.
- #ai
- #microsoft
- #openrouter
- #llm-verification
- #ai-agents