· via Hacker News – Front Page (native)
Strands Agents open-sources 2B decision model that runs locally in milliseconds
Strands Agents has released Strands Decider 2B, an open-source decision model that picks between options with confidence scores, answering in milliseconds on consumer hardware with weights and training data included.

The Strands Agents team has released Strands Decider 2B, a compact open-source model designed to make decisions rather than generate text. According to the announcement on the Strands Agents blog, the 2-billion-parameter model picks between offered options and assigns simple scores, and is tuned for quick experimentation and local development on hardware developers already own.
Deciding instead of generating
Strands Decider belongs to a category the blog calls decision models, also described as system one models, a class it says has drawn growing attention since TypeSafe AI launched its Jev model earlier this month. Instead of producing arbitrary output, these models answer constrained questions, such as whether a phrase refers to a coffee machine, which of several teams should handle a support request, or a sentiment score between 0 and 1.
The trade-offs are explicit. Decision models are faster and more capable at a given size, always return an answer from the selected options, and run at very low latency. In exchange, generating all outputs in a single parallel pass makes them significantly weaker than reasoning models on complex problems, and the inability to produce text rules out coding, chatbots, document summarization, and similar LLM tasks.
The blog highlights two additional properties: every decision carries a reliability score indicating how trustworthy the answer is, which frontier LLM inference APIs do not expose, and it is efficient to ask several questions about the same prompt.
Under the hood
The build starts from a pre-trained Qwen3.5-2B torso. The language-model head is removed, taking away text generation, and replaced with a pointer head of just over a million parameters. That head scores the hidden state at each option position against the hidden state at a special answer position, and the torso is fine-tuned with a rank-16 LoRA adapter.
The released model is the second major iteration of the architecture and the nineteenth version overall; an earlier variant used a slot head that performed significantly worse, and the repository documents what changed at each step. Weights are on Hugging Face, with code, training data, and training scripts on GitHub.
Benchmarks and latency
The team evaluates accuracy and calibration together on the public JevBench set, using the Brier score for calibration. Strands Decider 2B places third of 33 models in the 2B class, and first of 30 once just-over-2B models are excluded.
On latency, the blog claims answers to meaningful questions in tens of milliseconds, while measured results show a median of roughly 115 ms on a local Nvidia RTX 3090, rising approximately linearly with task size in tokens. On an M3 MacBook, small tasks come in at a median of about 153 ms. The model also completes 100 percent of JevBench's easy tasks, which the team maps to the routine decisions agents face in practice.
Why 2 billion parameters
The team cites two reasons for the size. One is experimentation: the model can be run and even trained on existing hardware, making trials fast and low risk. The other is a sweet spot, small enough to iterate on quickly but large enough to do meaningful work.
Where it fits
Reported uses for this class of model include model routing, tool selection, evaluation, guardrails, memory, context management, and policy classification. The blog also points to hybrid agents, where an LLM handles the hardest decisions and a decider model takes the easier, rote ones to cut cost and latency, plus experiments pairing deciders with fixed workflow languages and uses in games, task automation, and maze navigation.
Getting started is a single install: pip install strands-decider. The CLI can route a support complaint to a team, for example scoring billing at 0.845 against retail and sales. The repository also ships an example Strands agent that runs locally, uses a Bedrock LLM as its main model, and consults the decider as a guardrail: before an over-eager weather tool call fires, the decider answers two yes/no questions about whether the arguments are grounded in the conversation and whether it is too early to call the tool, sending the agent back to ask which city the user meant.
Why it matters
Agentic pipelines burn tokens and latency on decisions, such as routing, tool checks, and safety gates, that rarely need a frontier model. A 2-billion-parameter decider that answers in around a hundred milliseconds on consumer hardware, and attaches a calibration score to each answer, is a credible building block for that layer. The release is also unusually complete for its size: architecture, training data, scripts, and nineteen documented iterations are all public, so others can verify, fork, and improve it. If decision models become a standard component of agent stacks, an open reference implementation that anyone with a laptop can run and retrain lowers the barrier considerably.
- #open-source
- #agentic-ai
- #decision-models
- #local-llm
- #machine-learning