deniz.in

Markets

Weather

Loading weather

· via Cloudflare blog

Cloudflare open-sources Clef decision models and launches RL fine-tuning platform

Cloudflare has released Clef and Clef-flash, open-source decision models that return typed answers in milliseconds, plus an RL fine-tuning service to adapt them on Workers AI.

Cloudflare open-sources Clef decision models and launches RL fine-tuning platform

Cloudflare releases Clef and Clef-flash

Cloudflare has launched Clef, a family of decision models comprising a larger precision model and a faster variant called Clef-flash. Both are served on the company's Workers AI platform and published as open weights on Hugging Face under an Apache 2.0 license. Alongside them, Cloudflare announced a reinforcement-learning fine-tuning product for adapting Clef to specific workloads. According to the Cloudflare blog, Clef currently leads the Jev Decision Index, a benchmark for this category of model.

The release follows weeks of attention on decision models, a concept brought to prominence by Typesafe AI's Jev System One. Where a general LLM is open-ended and largely non-deterministic, a decision model returns typed answers with probabilities, fitting the moments in a workflow where software needs a dependable classification rather than generated text — and, unlike traditional classifiers, it can handle new categories without constant retraining.

What a decision model is for

The Cloudflare blog illustrates the pattern with a support-ticket example: pass in a customer message, ask whether it is urgent and which team should handle it, and the model replies with typed fields and confidence scores that code can act on to route, escalate or hand off to a human. Cloudflare says it has already been using Clef internally on its Threat Intelligence team to classify websites. Handed a domain, and using Browser Run to fetch and render the page, Clef returned probabilities such as 95% fashion, 85% ecommerce and under 1% phishing in 2.2 seconds. The company's fastest general LLM, gpt-oss-120b, took 4.7 seconds for the same workflow and produced only two classifications.

Clef separates itself from Jev in two ways, according to the post: it includes a vision encoder so it can classify images as well as text (Jev handles text only today), and it carries a 64k context window against Jev's 32k. Outputs are strictly typed and the API is Jev-compatible, which Cloudflare says makes swapping models straightforward.

Benchmark and latency claims

Cloudflare reports results across 43 evaluations. On the Jev Decision Index suite, Clef and Clef-flash score 98.47 and 98.76 on BFCL case-exact versus Jev's 95.75, and 79.60 versus 62.55 on the PhishNChips phishing set, though Jev still wins individual tests such as When2Call accuracy, at 80.97 against Clef's 72.37. On Typesafe's own eval suite, Cloudflare says its models beat Jev in three of four workflow areas, with Clef-flash singled out as especially strong given its speed.

Latency is the headline figure. Across those benchmarks, median response times were 209.3 ms for Clef and 38.8 ms for Clef-flash, compared with 524.1 ms for Jev, with p95 times of 238.6 ms and 122.4 ms respectively. Only Laya was faster, at a 5.8 ms median, but Cloudflare notes it trades away quality. Because the models run on Cloudflare's edge GPUs, the company argues they can sit in an agent's hot path, with an LLM on Workers AI then executing the chosen action.

Open weights and fine-tuning

The Apache 2.0 release means developers can run the models locally, while the hosted version comes with an enterprise guarantee that Cloudflare does not read, store, or use customer requests and responses for training — except for those opting into fine-tuning. That service begins as a hands-on engagement with Cloudflare's forward-deployed engineers, with a self-serve platform to follow; tuned models are redeployed onto Cloudflare. The blog also traces Clef's origins to experiments, published the week Jev launched, adapting the DiffusionGemma model to emit deterministic probabilities from its logprobs.

Why it matters

Decision models attack a real cost in AI engineering: using a large, non-deterministic LLM for simple routing decisions is slow and expensive. A model that answers typed questions in tens of milliseconds, can be fine-tuned with reinforcement learning, and ships under a permissive open-source licence gives teams a practical alternative — and the Jev-compatible API keeps switching costs low. With Typesafe, Cloudflare and others now competing on the same public benchmarks, structured decision-making looks set to become a distinct, commodity layer in agentic stacks.

  • #cloudflare
  • #decision-models
  • #open-source
  • #workers-ai
  • #reinforcement-learning

Related posts