· via dev.to (home feed)
WaterSheep is an open-source, API-compatible alternative to TypeSafe's Jev decision model
WaterSheep is a new open-source decision model that mirrors TypeSafe's pay-per-token Jev API, runs locally, and ships Apache 2.0 code, weights and a training pipeline.

An open rival to TypeSafe's Jev
A developer has released WaterSheep, an open-source model built as an alternative to TypeSafe's Jev for applications that want language models to make decisions rather than write text. Writing on dev.to under the handle samratduttaofficial, the project's author describes Jev's core idea: you submit a state and a set of typed questions, and instead of prose you get back options, scores and probabilities that your code can act on directly. Jev, however, is closed and metered per token, which is the gap WaterSheep is meant to fill.
Four question types, each with probabilities
WaterSheep answers questions in four formats. noul returns yes or no together with the probability of yes; choice picks exactly one option from a caller-supplied list; score places the input on a rating scale and reports the expected level; and multi returns every option that applies, each with its own probability. According to the post, the last of these is a capability Jev itself does not offer. Responses always include a probability for every option, not just the winner.
Built for drop-in compatibility
The compatibility story is the project's main selling point. WaterSheep runs as a local server: after installing it from GitHub and launching it with a serve command, it exposes a POST endpoint at /v1/systemone that accepts the same request and response shapes as Jev. Because the schemas match, TypeSafe's own Python SDK works against it — point the SDK at the local address and existing Jev scripts run unchanged. The author verified this with version 0.7.2 of typesafe-sdk; the JavaScript SDK has not been tested yet. Any API key value is accepted for local calls.
If you do not use an SDK, the endpoint also speaks plain HTTP, so a simple curl request is enough to query it. The model can additionally be loaded directly in Python through a standard transformers pipeline, with trust_remote_code enabled, and an ONNX build and a hosted demo are available for trying it without any local installation.
Reported accuracy and calibration
On the in-distribution test split, WaterSheep reaches 77.8% accuracy; on held-out datasets it never saw during training, that figure falls to 61.2%. The author also reports that the output probabilities are well calibrated, with expected calibration error of 0.026 and 0.043 in the two settings, meaning the model's stated confidence closely tracks how often it is right. That calibration matters operationally: confident answers can trigger automated actions while borderline cases are escalated to a human. Per-benchmark results, including the weak ones, are published in the project README, with further detail in an accompanying paper.
The stated limitations are equally plain. The model is English-only, long inputs get truncated, rating questions are its weakest format, and it is not intended to make high-stakes decisions on its own.
Open weights, open training
Code, model weights and the training pipeline are all Apache 2.0 licensed, and the whole thing can be retrained on custom data with a single script, according to the post. The project is independent and not affiliated with TypeSafe. The author is openly asking for feedback, particularly examples of inputs where the model gets things wrong, and the model is available on GitHub and Hugging Face.
Why it matters
Decision-only models sit in a different part of the stack from chat assistants: they replace hand-written rules or chained prompts in triage, routing and classification pipelines, where per-token billing on a closed API adds up and sends data to a third party. WaterSheep changes the economics by letting teams run the same workload locally, on their own hardware, with no vendor lock-in. The Jev-compatible API keeps switching costs low — existing code can be pointed at a local server for evaluation before anything is rewritten.
The honest reporting cuts both ways. A 61.2% accuracy on unseen datasets means this is a starting point and a fine-tuning base rather than a wholesale replacement, and the author says as much. But with calibrated probabilities, a question type Jev lacks, and an Apache 2.0 training pipeline that invites retraining on domain data, it gives developers a credible open reference implementation of the decision-model pattern — and leverage they did not have before.
- #open-source
- #machine-learning
- #llm
- #self-hosting
- #apache-2-0