deniz.in

Markets

Weather

Loading weather

· via Vercel blog

Vercel adds Jev classifier support to Python AI SDK via experimental evaluate API

Vercel's AI SDK for Python now ships an experimental evaluate() API for Jev, a new model class that answers structured multiple-choice questions with confidence scores instead of generating text.

Vercel adds Jev classifier support to Python AI SDK via experimental evaluate API

What Jev is

A new category of AI model called Jev is drawing wide experimentation, according to a post on the Vercel blog, with people applying it to everything from trading decisions to generating user interfaces. Unlike a language model that produces free-form text, Jev behaves as a universal classifier: you hand it some data and a set of multiple-choice questions, and it replies with answers plus a confidence score for each.

Vercel explains that classic classifiers — a spam filter, say — require labelled data and task-specific training for every new job. The team behind Jev instead worked out how to convert an LLM into a classifier that needs no such training, because it already carries broad general knowledge. The result is cheap and fast: rather than generating prose, it returns structured JSON that conforms to the types and choices you define in your questions. It still makes mistakes, Vercel cautions, and whether it is accurate enough is something each use case has to establish through testing.

The evaluate() API in Python

To make the model easy to experiment with, Vercel's Python team has released a version of the AI SDK for Python containing an experimental evaluate() function that connects straight to the model. Setup amounts to installing the package with uv add ai, creating an AI Gateway key, and setting AI_GATEWAY_API_KEY; calls target the typesafe-ai/jev model.

Vercel says the Python surface mirrors Jev's official API almost one-to-one. There is a single function, evaluate(), which takes a model, a state (either a plain string or arbitrary JSON), and a mapping of questions. Questions come in three shapes: ChoiceQuestion for picking a single option from a list, ScoreQuestion for rating the input against a scale you define, and NoulQuestion for estimating how probable a given statement is.

Two experiments, mixed results

The post then walks through two examples that the author openly grades as uneven, inviting readers to decide which is which.

The first uses Jev literally as a classifier: deciding whether text typed into an agentic Python REPL is English or Python code. The author had previously spent two days training a custom classifier for this and abandoned it, because the UI kept flipping between languages mid-word. Jev performs noticeably better, though gaps remain — it reads the partially typed expression "what's" + " up" as English, where a person would recognise incomplete Python.

The second experiment pushes Jev into work it was never meant for: writing code. Because Jev cannot produce text, the author first tried having it pick characters one at a time, then pivoted to assembling an abstract syntax tree through a sequence of choices. An LLM — GPT-5.6, in this case — first expands the user's prompt into detailed instructions, since without a plan Jev struggles with even elementary tasks. At each step Jev sees the program so far, the field being filled, and candidate nodes with a preview of the code each would produce, while the surrounding program handles punctuation and indentation. The output was syntactically valid but mostly incorrect Python, and the author's takeaway is blunt: making classifiers do the job of a generative model is hard. The code is published on GitHub for anyone who thinks they can do better.

Why it matters

Jev represents a different bargain from chat-style models: no prose, just fast, cheap, typed decisions with attached confidence. That profile suits a range of production problems — routing, reranking, filtering, thresholding — where teams currently pay full LLM prices and then wrestle with parsing output. Shipping it as a first-class SDK call lowers the barrier for Python developers to test that idea, and the post's candid account of where Jev falls short is useful calibration in itself. As Vercel frames it, this model class is only starting to be explored, and the honest verdict for now is that it earns its keep on narrow decisions rather than generation.

  • #ai
  • #python
  • #sdk
  • #machine-learning
  • #vercel

Related posts