deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

TypeSafe's Jev classifier barely beats a flat guess in physics calibration test

An independent analysis tested TypeSafe's System One classifier Jev against ten physics distributions with known answers and found its probabilities barely better, and sometimes worse, than an even spread.

TypeSafe's Jev classifier barely beats a flat guess in physics calibration test

What Jev is

TypeSafe has released Jev, a "System One" classifier positioned next to large language models rather than against them. According to an independent analysis published on the Maximum Effort Substack, which reached Hacker News' front page, the name nods to Daniel Kahneman's Thinking, Fast and Slow: LLMs such as Claude or GPT play the role of slow, deliberate System Two reasoning, while "Decision Models" like Jev and its predecessor Laya are built for fast, instinctive System One judgments.

The analysis sketches the architecture: a pretrained transformer keeps its generality, and a classifier is attached to its output. Instead of freeform text, a caller supplies context and a multiple-choice question and receives a typed answer with a probability distribution across the options. The author contrasts this with a typical LLM call, where asking a chat model what 2 + 2 equals returns a string you still have to parse, while Jev returns the integer 4 directly. The entire evaluation reportedly cost under $4.00 to run.

A separate tutorial on dev.to shows the intended production pattern. TypeSafe's TypeScript SDK takes one shared state object plus a set of independent questions, using primitives named noul, choice and score, all evaluated in parallel against that state. Questions point at data through dot-and-index paths such as ticket.messages[0].text, and the tutorial advises shallow schemas, immutable updates for cheap diffing, and isolating untrusted text at the schema level to blunt prompt injection. The piece promotes the author's paid ebook, so its performance claims deserve corresponding skepticism.

Testing calibration against physics

The Substack author's question was simple: when Jev reports a probability, does it mean anything? The testing ground was questions with exactly known answers, namely physics probability distributions. Ten families were used (Gaussian, Lorentzian, Maxwell, Gamma, Exponential, Rayleigh, Uniform, Poisson, Binomial and Boltzmann), each with five prompt templates and 20 variations per template, for 1,000 settings in total. Jev was asked to choose among fixed bins of a continuous parameter, so a well-calibrated model should assign each bin the probability mass implied by the true distribution. GPT-6 Astra and Claude Opus 5.5 wrote the templates, API calls and plots.

The results

Scoring used total variation distance: 0 means Jev's distribution matches theory, 1 means completely disjoint. Jev averaged 0.518. Had it abandoned any claim to knowledge and divided its probability uniformly across all choices, it would have averaged 0.546, barely worse. On Uniform questions Jev scored 0.77 against 0.39 for that uninformed baseline, so an even spread was dramatically better calibrated than the model; on Poisson the two were effectively tied at 0.65 versus 0.64.

Two failure patterns stand out in the analysis. First, Jev's output distributions are overly concentrated, and it mishandles smoothly vanishing tails by placing nonzero weight on bins where the true probability is nearly zero. The Lorentzian scored relatively well, which the author suspects is because that family is inherently sharp-peaked and heavy-tailed. Second, apparent competence at locating a distribution's peak may partly be copying: in Gaussian-style prompts the center is stated in the prompt itself. For Maxwell, Rayleigh and Gamma, where the peak must be derived rather than read off, Jev found it in only about 20% of settings and returned very flat distributions, and for Rayleigh and Gamma it produced plausible-looking uniform answers, which the author interprets as the model signalling uncertainty.

Consequences

The author flags an immediate implication: evaluation frameworks such as JevEval, which lean on these models as automated judges of LLM answers, look shaky when the judge's own probabilities miss badly on textbook problems. That concern lands directly on the usage pattern the dev.to tutorial demonstrates, where business logic fires only when several Jev probabilities all exceed 0.8. Thresholding on a score only works if the score is calibrated.

Why it matters

Jev represents an emerging class of LLM-adjacent decision models promising cheap, typed, probabilistic classification in fast response loops instead of text generation. That promise rests entirely on the probabilities being trustworthy, and a sub-$4 experiment suggests they currently are not; on some families the model did worse than knowing nothing and guessing flatly. Teams wiring such scores into automated decisions should treat them as weak ranking signals at best until calibration improves, and the experiment itself is a reusable, low-cost template for auditing any model that claims to output calibrated probabilities.

  • #machine-learning
  • #calibration
  • #classification
  • #llm
  • #evaluation

Related posts