deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Jeff fine-tunes small open models into ~30 ms decision classifiers compatible with Jev

An independent project has fine-tuned Qwen3.5 and Gemma 4 into tiny decision-only classifiers that score options in roughly 30 ms on consumer hardware, edging past Jev's published results on classification benchmarks.

Jeff fine-tunes small open models into ~30 ms decision classifiers compatible with Jev

An independent project called Jeff has released three small fine-tuned models that do exactly one job: take a described situation plus a list of plain-language options, and return a calibrated probability for each option in a single forward pass. According to the project's GitHub repository, decisions take about 22 ms on an NVIDIA RTX PRO 6000 and 28 ms on an Apple M4 Max running MLX. The models accept the same request format as Jev, the decision-only model, though the project states plainly that it is not affiliated with or endorsed by Jev's maker, TypeSafe.

What the models do

The release covers three checkpoints hosted on Hugging Face: Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B. No text is generated and nothing needs parsing — a request describes a state, such as a customer refund message, and asks questions about it. Three question types are supported: choice (picking one of up to 255 options), noul (a yes/no probability) and score (a point on a user-defined scale). Several independent questions can be answered together in one request.

Classification is zero-shot, so categories like support queues, user intents, moderation labels or voice commands do not need to appear in the training data; they are described in words at request time. The repository's usage guidance is to keep planning logic in code and treat Jeff purely as a decision-maker: when asked to forecast outcomes rather than choose between stated options, it performs no better than random.

Strong on classification, far behind on reasoning

The project evaluated 4,599 questions across five public benchmarks, plus JevBench's 105-item hard tier scored separately. Jeff-Qwen3.5-2B reached 83.1 overall, just above Jev's published 83.0, with AutoJev-27B at 84.9 shown for reference. Jeff clearly beat Jev on Financial PhraseBank (96.3 versus 77.0) and RAGTruth (88.9 versus 77.3), and lost on WinoGrande (79.0 versus 90.7).

On reasoning-heavy sets the gap widens sharply: BBH (68.0 versus 94.3), JudgeBench (64.6 versus 78.6) and JevBench hard (53.3 versus 73.3). The repository also cautions that the published Jev and AutoJev figures were measured on a different sample of the same benchmarks, so the comparisons are indicative rather than controlled.

Games as an unplanned zero-shot test

To probe tasks unlike the benchmarks, the authors had the models play Doom, Frogger and Pac-Man, 20 episodes each, with the game state and legal moves described in words. The 0.8B model tied a hand-coded rule bot on Doom kills (both 6.55) and nearly matched it on Frogger crossings (10.3 versus 10.25), while collecting 57.0 of 98 Pac-Man pellets against the bot's 94.1. Jeff decided in 29–49 ms per move on the M4 Max; Jev's published Doom run took 212 ms per API call, though the project notes the two timings were not measured on the same hardware. Oddly, the larger 2B model played worse than the 0.8B despite higher benchmark scores, and the authors recommend the 0.8B as the best fit for fast option picking.

Trained at home, fine-tunable in half an hour

Everything was built on local hardware: the 0.8B trained in roughly two hours on a single RTX PRO 6000 workstation GPU and the 2B in about 3.5 hours. All synthetic training data was written by the open model Qwen3.8-Flash-Next on two DGX Sparks, with a closed model used only to spot-check a sample of that data. The training code builds on the open-source AutoJev recipe. The weights are small — 1.7 GB for the 0.8B and 4.2 GB for the 2B in 16-bit.

Fine-tuning is where the practical payoff sits. A voice-navigation fine-tune on roughly 11,000 app-specific examples took about half an hour on one GPU and lifted held-out accuracy from 31.7% to 95.8%, at around 40 ms per decision on the M4 Max.

Why it matters

A separate dev.to essay about Jev makes the economic argument: most enterprise agent traffic is routing, triage, classification and gating rather than prose generation, and a single scoring pass replaces many seconds of token generation with a fraction of a second of computation. Jeff pushes that logic onto local consumer hardware. At roughly 30 ms and under 2 GB of weights, high-volume routing decisions can leave the API billing path entirely, and the demonstrated half-hour fine-tune turns a weak zero-shot performer into a production-grade classifier. The limits are equally clear: on reasoning-heavy work the small models remain far behind large ones, so the pattern that emerges is a division of labour — tiny calibrated deciders in the hot loop, big models reserved for the hard thinking.

  • #open-source
  • #machine-learning
  • #llm
  • #classification
  • #fine-tuning

Related posts