deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Cloudflare's Clef open-weights decision model nearly matches paid hosted API in local benchmark

A developer benchmark on dev.to pits Cloudflare's new open-weights decision model Clef against the hosted Jev API it aims to replace, and finds the free 9B variant matches the paid service on the highest-stakes agent decisions.

Cloudflare's Clef open-weights decision model nearly matches paid hosted API in local benchmark

Cloudflare has released Clef, an open-weights model built to make decisions rather than generate text. According to a hands-on benchmark published on dev.to on October 1, the smaller Clef-Flash 9B variant, running locally on a single RTX 3090, matched a paid hosted API on the agent decisions with the highest stakes while trailing it on nuanced classification.

What Clef is

Per the dev.to post, Clef ships in two sizes, a 27B flagship and a 9B Flash variant, both built on Qwen backbones and equipped with a joint schema head that outputs calibrated probabilities over typed questions. Instead of producing prose, developers hand the model a state blob plus choice questions (with criteria per option) or binary questions, and get back a probability for each option from a single forward pass.

The release targets agent builders who currently pay per call for TypeSafe's Jev API, which Clef is positioned to replace outright. The author describes the common alternative it displaces: wiring a chat model plus regex parsing into message routing, GUI action selection or long-running job supervision, then fighting prompt drift between calls.

The benchmark setup

The author ran Clef-Flash 9B in bf16 under PyTorch against Jev-1.13.0 over its hosted API, using identical prompts, typed questions and option criteria. The workload was not synthetic: 42 labeled decisions taken from a real autonomous coding agent that works overnight hunting paid GitHub bounties. Three families were tested. Ten computer-use action choices drawn from prevalidated action tables, including two deliberate safety traps where the correct answer is to do nothing. Twelve subagent supervision calls with a five-way classification covering progressing, waiting on an answer, stuck in a loop, blocked and finished. Twenty inbox-style triage routes labeled now, today, queue or ignore.

Results

The headline score favors the incumbent: 71.4% overall accuracy for Jev against 66.7% for Clef-Flash. But the per-task split is the real story, and it flatters the local model.

On computer-use action choice both models scored a perfect 10 for 10, correctly abstaining on an irreversible money transfer with missing details and choosing to reobserve after a spinner that might have swallowed a duplicate click. On supervision both went 8 for 12 and missed the same two cases; one of those the author concedes is ambiguous enough to dispute their own gold label. Triage was the only genuine gap, with Jev at 12/20 versus Clef's 10/20. The hosted model won mostly on today-versus-queue boundary calls, and the author notes that a 9B model losing a nuance contest to a much larger hosted one is expected.

Latency was close: Jev posted a p50 of 225 milliseconds over the network, while Clef-Flash delivered 315 milliseconds fully local with no batching. The two models agreed on 35 of 42 predictions, roughly 83%, and where they diverged they tended to fail for the same underlying reasons.

Running it locally

Cost is the obvious motivation, but the author argues privacy is the stronger one: supervision calls carry log tails, file paths and message contents that would otherwise leave the local network, and a local model keeps functioning without internet access. The upgrade path is hardware-bound. The full 27B model reportedly beats Flash on exactly the nuance tasks where Flash trails Jev, but needs about 54GB of VRAM in bf16 against a 3090's 24GB, so better local decisions effectively mean buying a second GPU.

One deployment caveat stands out: community GGUF quantizations on Hugging Face cover only the language backbone. The joint decision head requires the transformers code path, so the decision functionality cannot run in llama.cpp today. A full local setup means PyTorch and the custom joint schema model file from the repository.

The author's resulting arrangement keeps the 9B model local as the default engine for supervision and action-choice calls, free and fast enough, while triage stays on the hosted API because nuance there justifies the per-call cost.

Why it matters

Clef marks a deliberate departure from the habit of stretching general-purpose chat models into every role. Routing, gating, ranking and action selection are decision problems, and a purpose-built model with calibrated probabilities removes the output parsing and drift that make chat-model agents brittle. The benchmark also shows what open weights change in practice: with latency roughly tied, the hosted-versus-local choice reduces to where your data lives and what each call costs, and local wins both for sensitive workloads.

Caveats apply. This is a 42-case evaluation from one person's agent, not a paper, and the triage labels involve judgment calls where reasonable operators route differently. Still, the author's takeaway generalizes well: running your own evaluation tells you less about which model is best than about which of your calls actually need a paid API at all.

  • #cloudflare
  • #open-weights
  • #ai-agents
  • #local-llm
  • #benchmarks

Related posts