deniz.in

Markets

Weather

Loading weather

· via Cloudflare blog

Cloudflare ships multimodal Clef-omni model and cuts Clef-flash pricing below Jev

Cloudflare's Clef-omni takes text, images, audio and video in a single open-weight decision model call, while Clef-flash pricing drops to $0.038 per million input tokens.

Cloudflare ships multimodal Clef-omni model and cuts Clef-flash pricing below Jev

Cloudflare has expanded its Clef family of open-weight decision models with Clef-omni, a version that accepts audio, video, image and text input in a single API call. Announced on the Cloudflare blog alongside a price cut for Clef-flash and serving-layer speedups for Clef, the release lands roughly a week after the original Clef and Clef-flash launch. According to the company, the first Clef went from a Friday-evening decision to a Thursday launch, with training completed over a single weekend.

One call instead of a pipeline

Decision models, which return structured schema-bound judgments rather than generated prose, have been mostly text-only since TypeSafe released Jev, Cloudflare notes. The original Clef added images and video frame arrays; Clef-omni adds audio in wav or mp3 form and full video in mp4 or webm. The practical appeal is less plumbing: instead of chaining a speech-to-text model or extracting audio and image channels from a video before analysis, one request scores everything. Cloudflare's example sends a photo, an audio recording and a video of an appliance installation together, asking separate questions about the serial label, how the unit sounds and whether the fan is running.

Built on Qwen3-Omni without the speech output

Clef-omni sits on the Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation, which already handled text, imagery, audio and video in one pipeline. Cloudflare kept the comprehension backbone and discarded the text-to-speech output components. Because Clef models are not LLMs and skip output token generation, incoming media maps directly into a unified sequence with audio and video synced to visual frames. Candidate values are pulled from internal embeddings through two-stage attention routing, and a built-in lexical grammar preserves option semantics for fast schema-constrained scoring. Training mirrors Clef: the Qwen3 backbone is frozen, low-rank adapters are trained, and label-smoothed cross-entropy loss is combined with Brier score calibration.

Reported latency is a median of about 130 ms for text-only decisions, 150 ms with images, a few hundred milliseconds for audio clips, and around 1.5 seconds for a 21-second video clip with sound, all in one call.

Benchmarks: strong in places, not a clean sweep

Cloudflare published results across its own suite and TypeSafe's evals, and no single model dominates. Clef-omni scores 97.7 macro-F1 on CLINC150+OOS against Jev's 89.27, and 98.2 case-exact on BFCL, but trails Jev on When2Call accuracy at 63.3 versus 80.97 and on BRIGHT nDCG@10 at 42.0 versus 47.52. Clef-flash posts the best home appliances result at 97.73 case-exact, while Clef leads ToolRet at 69.1 nDCG@10. On TypeSafe's evals the picture splits: Clef leads invoice processing exact actions with 64.75 against Jev's 61.8, the Clef variants lead customer service, security incidents are essentially tied, and Jev keeps an edge on agent trace observability at 71.6 versus Clef-omni's 65.8.

Cheaper Clef-flash with a smaller hosted context

Clef-flash drops from $0.09 to $0.038 per million input tokens, which Cloudflare says makes it cheaper than Jev. Clef holds at $0.24 per million, and Clef-omni launches at $0.15. The discount carries a catch: the hosted Clef-flash context window shrinks from the previously advertised 64k to 24k tokens. Cloudflare's usage data shows only 0.24% of requests exceed 24k input tokens, so the company made the cut to reach the lower price. The weights on Hugging Face are untouched and support a 256k context when self-hosted, and Clef itself keeps its 64k window for larger payloads.

Clef serves faster

The hosted Clef model on Workers AI also got faster through serving-infrastructure changes rather than new weights. For roughly 800-token inputs, median latency fell from 262 ms to 152 ms, with p95 improving from 438 ms to 351 ms.

Why it matters

Two things stand out. First, pace: Cloudflare turned an idea into shipped open-weight models in under a week and has already iterated twice, with weights available on Hugging Face for anyone to self-host. Second, economics: at $0.038 per million input tokens, a decision model becomes cheap enough to embed as a routine component across agentic workflows, and single-call multimodality removes the transcription and demuxing steps that usually precede analysis. The trade-offs are equally instructive. A trimmed hosted context window set against unchanged open weights shows where managed convenience and self-hosted flexibility diverge, and the mixed benchmark results suggest the decision-model category is still young enough that no offering, Cloudflare's included, is a default choice yet.

  • #cloudflare
  • #open-weights
  • #multimodal
  • #decision-models
  • #ai

Related posts