# Cloudflare open-sources Clef decision models at 209.3 ms versus Jev's 524.1 ms

From The Forward Pass daily issue, October 2, 2026 (https://theforwardpass.net/archive/daily/2026-10-02). Source: https://theforwardpass.net/archive/daily/2026-10-02/cloudflare-open-sources-clef-decision-models-at-209-3-ms-versus-jevs-524-1-ms

> This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.

**Top News** · 472 HN points

Some LLM calls exist only to pick an option, like classifying a website. Cloudflare's Clef and Clef-flash are decision models built for that job. They classify inputs and return structured answers with probabilities your software can act on.

Here is what makes this credible:
- Across 43 benchmarks, Clef's median latency is 209.3 ms and Clef-flash's is 38.8 ms, versus Jev at 524.1 ms.
- On BFCL case exact, Clef scores 98.47 and Clef-flash 98.76, versus Jev at 95.75.
- Inference runs one prefill-only pass, then scores valid schema choices in parallel instead of generating text.
- Clef adds image input and a 64k context window, where Jev is text-only at 32k.
- The backbones are Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, and the API is Jev-compatible.

Fine-tuning starts as a hands-on service with Cloudflare engineers, with self-serve planned later.

One catch: Jev still leads on When2Call accuracy (80.97 versus 72.37) and BRIGHT nDCG@10.

Try it: POST to `https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef` with a bearer token, or download the Apache 2.0 weights from Hugging Face.

Sources: [blog.cloudflare.com](https://blog.cloudflare.com/clef-decision-models/)
