deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Microsoft-Decision-1 scores options for agent routing, lands in Foundry at $0.042 per million tokens

Microsoft released Microsoft-Decision-1, a small model that returns a calibrated probability for each option in a list, built for routing and agent decisions and priced at $0.042 per million input tokens.

Microsoft-Decision-1 scores options for agent routing, lands in Foundry at $0.042 per million tokens

A scoring model rather than a chat model

Microsoft released Microsoft-Decision-1 on 9 October 2026, a compact language model built for a single task: hand it a situation and a fixed list of options, and it returns a calibrated probability score for every option on that list. According to a dev.to write-up of the vendor's launch post on its Command Line publication, the model is available in Microsoft Foundry from day one, with access through OpenRouter planned.

The model was created by post-training Qwen3.5-9B for one-pass decision scoring, and Microsoft says it will rebuild it on top of other base systems, including its own Microsoft AI models and OpenAI models. Achint Srivastava, a Microsoft vice president in the Office of the CTO, drew a line between this class of model and generative ones: where a chat model produces text or works through open-ended problems, a decision model exists to return structured output that software can act on immediately.

That shape defines the use cases. Routing a support ticket, checking a proposed tool call against a rubric, and grading an AI response all reduce to the same call: a context, a closed list, and a score for each entry. The calling program keeps control of the cutoff; the model only ranks the options.

The numbers, all from the vendor

Microsoft reports the highest accuracy in a comparison across 36 benchmarks covering close to 150,000 questions, with tests spanning routing, ranking, long context, multilingual and out-of-distribution tasks. It also claims P50 latency roughly 35 times faster than GPT-6 Sol. The dev.to piece is explicit that the performance and pricing figures come from Microsoft's own benchmarking rather than independent review.

Pricing, recorded by Windows Report, starts at $0.042 per million input tokens, with output tokens free. At that rate, a support desk scoring 200,000 inbound messages a month pays under nine dollars for input, which turns a check teams often skip for cost reasons into one that can run on every message.

Microsoft's latency argument is simple arithmetic: adding 100 milliseconds to each of 20 sequential decisions adds two seconds to a workflow. Agent pipelines with dependent steps inherit that delay at every hop.

Five problems, two with hard numbers

Microsoft names five engineering problems it had to solve: speed, quality that holds up on unseen tasks, robustness to reworded inputs, probability calibration, and safety. Robustness comes with a figure: the model changes its decision on 1.3 percent of perturbations on average across eight ways of rewriting a request, with zero flips when option descriptions are paraphrased or when the options are reversed or shuffled. Calibration is framed in operational terms — a 90 percent prediction should be right about nine times in ten on representative cases, because applications use confidence to decide whether to act, defer, or escalate.

Where Microsoft already runs it

The launch post describes four internal deployments. Xbox Research used the model to sort more than 10,000 open-ended feedback items from surveys, Steam and Twitter/X into themes researchers had fixed in advance; Microsoft reports the result as competitive in quality with GPT-6 Sol while running over 14 times faster and 200 times cheaper. The Copilot team, measuring chat and agentic response quality, found it competitive with GPT5.6 Luna and 100 times faster. On-call engineers use it for knowledge retrieval during live incidents. Microsoft Discovery uses it to grade experiments inside an adaptive replanning loop, reported as 46 times more consistent than an LLM-based scorer at three times the speed.

A shared pattern stands out: in every case the label set existed before the model did. Humans decided the categories or the rubric; the model performs the sort.

Why it matters

Microsoft-Decision-1 formalises a pattern many production agent stacks already use informally: a cheap judge sitting next to an expensive writer. The large model drafts; the small model answers one bounded question about what happens next. A dedicated scorer strips generated text out of a step where nobody reads the output, and the pricing makes the check affordable at message volume.

The caveats are as clear as the pitch. Every figure is vendor-reported, and Microsoft has published no deployment procedure, so evaluation stays with the buyer: a labelled set of your own tickets, a baseline to beat, and a view of which errors cost money, since a wrong proceed and a wrong escalate are not priced the same. If the claims hold up outside Microsoft's own tests, this points to the model layer splitting into specialised tiers — generation, reasoning and decision — rather than one model being asked to do all three.

  • #microsoft
  • #ai-agents
  • #language-models
  • #microsoft-foundry
  • #machine-learning

Related posts