Meta description (draft): Cloudflare open-sourced Clef: a 27B scoring head on a frozen Qwen3.8-27B that scores answer sets, takes images, and returns probabilities in 209 ms. —
Cloudflare open-sourced two decision models on October 1: Clef and Clef-flash, Apache-2.0 on Hugging Face, with weights on the Hugging Face repos and code paths documented in Workers AI. Clef is a scoring head and rank-256 LoRA on a frozen Qwen3.8-27B backbone; Clef-flash rides a frozen Qwen3.5-9B. Both return probabilities for fixed answer sets instead of generated text, and both accept images, which no other decision model in this lane does today. Independent aggregators noticed within a day; the table that circulated on X (Mia’s screenshot, 534 likes) matches the blog’s own benchmark table number for number across all ten rows, so the figures below carry the blog’s data with the caveat attached.
On the vendor’s own Decision Index leaderboard (marked self-reported, and The Register notes independent verification is still pending), Clef currently tops the class at 61.21 against Jev’s 57.91. Latency is the stronger claim for daily use: 209 ms median for the 27B model and 38.8 ms for the flash, against 524 ms for Jev on the same harness. Cloudflare’s own developer account frames the comparison more bluntly: “Like Jev? You’ll love Clef and Clef-flash, which are smarter, faster, and fully Jev-API compatible” (@CloudflareDev), and the vendors’ numbers behind it are the two lines above: latency 2.5x better against the full model and 13.5x against the flash, index accuracy 3.3 points, with the outlier rows wider (Home appliances 82.95 versus 52.27).
How it works
The mechanics are close to Kev’s approach with a multimodal twist. The frozen Qwen backbone does a prefill-only pass; nothing autoregressive happens. A separately trained routing head then scores every valid option of every schema-defined question in parallel, up to 64 questions per request. The routing head is a two-stage attention design: option-specific evidence extraction followed by joint cross-field attention, with a lexical prior that keeps intent anchored when you swap option phrasings.
Clef-flash keeps the vision encoder from its backbone: PNG, JPEG, WebP, up to 4 images per request, each capped at 4 MiB and 16 MP, with 8 MiB total decoded per request. Training used label-smoothed cross-entropy for the schema outputs plus Brier loss for probability calibration, and a secondary objective the Cloudflare team calls RLCD, Reinforcement Learning for Calibrated Decisions: partial credit for adjacent ordinal picks, a reward shape for fully precise outputs, and a reference-model penalty to stop distribution drift.
Self-hosting limits are concrete: the Cloudflare PM told The Register the flash needs at least 41 GB VRAM and the full model at least 85 GB, at single concurrency and 64k context. That puts the full model on the workstation tier (96 GB boxes) and the flash on enthusiast hardware; the hosted API is $0.24 per million input tokens for Clef and $0.09 for Clef-flash.
Strands Decider 2B: the AWS entry
One day earlier, AWS’s Strands Labs shipped Strands Decider 2B, and the origin story shows where this category is going: Marc Brooker saw TypeSafe’s Jev and built his own, then Amazon shipped it. Strands Decider 2B is a rank-16 LoRA and pointer head (about a million parameters of fresh structure) on a frozen Qwen3.5-2B base, with the training scripts and datasets published in the repo. On JevBench it ranks 3rd of 33 models in its size class and first of 30 excluding the just-over-2B entrants, with a 72.29 percent public-set accuracy in the model card’s own eval table.
The week’s shape
One week after Jev launched, the lane has a TypeSafe frontier, an open Apache-2.0 challenger from one CDN with images and a 64k window, and a 2B entry from a hyperscaler that runs on a Pi-class box. Ollama meanwhile ships the Jev API locally in 0.35 with Nimble and Tev1 pulls. Every one of these is a bet that most agent decisions are classification, not generation: model routing, tool selection, guardrails, triage, and evaluations all reduce to structured choices with confidence scores attached. The local case is not price anymore; it is the sub-100-ms loop and owning the decision boundary inside your own harness.
Open weights win this round on fit math: Clef runs a 55 GB checkpoint for the price of a 96 GB box, Clef-flash runs where Qwen3.5-9B runs, Strands Decider 2B runs on a CPU. If the Decision Index scores hold under independent verification, the classification workloads that quietly make up most of an agent’s bill stop being a frontier-model cost and start being a local line item.
Sources: Cloudflare blog - Workers AI clef docs - Cloudflare/clef on HF - Cloudflare/clef-flash on HF - Strands Decider 2B blog - StrandsAgents Decider on HF - The Register on the Clef launch - Cloudflare Developers’ Clef quote-tweet
Related on this site: Kev: an open decision-model family you can train yourself - Ollama 0.35 ships the Jev API locally - Laya answers yes/no questions in 33 ms - What is Jev? The plain-English guide
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.