THE FORWARD PASS
Editorial

Deep dive The Forward Pass research desk

Cloudflare's Clef replaces text generation with schema scoring for agent decisions

Cloudflare has released Clef and Clef-flash, open-weight decision models on Workers AI that return structured classifications with probabilities, and it reports latency well below Jev's.

AI-generated article

This article is researched and written by AI models from the sources it cites, and every claim is checked against them. No human edits it before it is published.

Kind
Deep dive
By
The Forward Pass research desk
Published
October 3, 2026
Read time
4 min

What shipped

Cloudflare released two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI and reachable through the Cloudflare API. The weights are also on Hugging Face under Apache 2.0, which Cloudflare says is "for you to run locally and experiment with yourselves."

Cloudflare's definition: "A decision model makes classifications to help agents decide how to act, based on certain probabilities." The target is the class of LLM calls that exist only to pick an option, such as classifying a website. The models produce structured classifications and probabilities that software agents can use to choose actions.

  • Clef is the larger model. Cloudflare says it "is currently the leader when evaluated against the Jev Decision Index."
  • Clef-flash is the faster model, at a reported 38.8 ms median latency.
  • The API is Jev-compatible, according to Cloudflare.

How it works

Both models sit on frozen Qwen backbones. Clef uses Qwen3.8-27B and Clef-flash uses Qwen3.5-9B. Cloudflare says it "jointly optimized the routing head alongside rank-256 low-rank adapters."

Inference does not generate intermediate text token by token. The model runs one prefill-only pass, then scores valid schema choices in parallel. Cloudflare describes a two-stage attention routing process: "every valid choice extracts context relevant to the prompt, allowing individual field parameters to cross-attend with other fields and back to the original payload prior to scoring."

Two capability additions over Jev, per Cloudflare: a vision encoder for classifying images (Jev is text-only today) and a 64k context window versus Jev's 32k, which Cloudflare says lets users "squeeze more input state for the model to classify against."

How it compares

Cloudflare reports median latency across the 43 eval benchmarks it ran:

| Model | Median latency (ms) | |---|---| | Clef | 209.3 | | Clef-flash | 38.8 | | Jev | 524.1 | | DiffusionGemma Jev | 84.4 | | Kev-9B | 51.4 | | Laya | 5.8 |

Cloudflare says Clef models beat the other decision models on latency except Laya, which it describes as very fast but trading off benchmark quality.

On accuracy, Cloudflare reports:

  • BFCL case exact: Clef 98.47, Clef-flash 98.76, Jev 95.75.
  • When2Call accuracy: Jev leads at 80.97 versus 72.37.
  • BRIGHT nDCG@10: Jev also leads.
  • Typesafe's eval suite: Cloudflare says its models beat Jev in 3 of 4 areas and that Clef-flash "performs exceptionally well, given how much faster it is."

Cloudflare also cites one of its threat-intelligence workflows. Clef fetched, rendered and classified a website in 2.2 seconds. gpt-oss-120b, which Cloudflare calls its fastest general LLM, took 4.7 seconds in the same workflow and returned only two classifications. That figure covers the full pipeline, including fetch and render, not model latency alone.

Access, privacy and fine-tuning

Hosted access runs through Workers AI at @cf/cloudflare/clef. Cloudflare calls the hosted models enterprise-ready and says it does not read, store or train on requests or responses unless the customer opts into its fine-tuning product.

Fine-tuning starts as a hands-on service with Cloudflare's forward-deployed engineer (FDE) team. A self-serve platform for training and redeploying onto Cloudflare is described as coming later, with no date given. Cloudflare notes the tradeoff: "When you fine-tune a model, you may give up some general purpose performance in exchange for higher accuracy in a specific domain."

Limits and open questions

  • The benchmark numbers are Cloudflare's. The latency and accuracy comparisons come from Cloudflare's own runs and should be read as vendor claims.
  • Jev still wins some tasks. The When2Call gap (80.97 versus 72.37) matters if your decisions resemble "should the agent call a tool now." Jev also leads on BRIGHT nDCG@10.
  • Only medians are reported here. Tail latency and behavior on long 64k inputs or images are not quantified in these figures.
  • Fine-tuning has a cost in generality. Cloudflare says tuning may reduce general-purpose performance.
  • Self-serve tuning has no date. Teams needing custom tuning today go through the FDE engagement.

Who should care

  • Agent builders whose pipelines spend LLM calls on routing, triage or yes/no gates. Clef returns structured answers with probabilities directly.
  • Teams already on Jev, since Cloudflare says the API is compatible and reports lower latency at higher BFCL scores.
  • Anyone classifying images or long inputs, given the vision encoder and 64k window.
  • Teams that want to run models locally, since the Apache 2.0 weights allow it. Clef-flash on a 9B backbone is the lighter option.

Try it

Call the hosted Clef model through the Workers AI API with Cloudflare's example request, which asks an urgency, routing and severity question about one support message. For local runs, download the Apache 2.0 weights from Hugging Face. For workload-specific tuning, contact Cloudflare's FDE team.

1. Classify a support request with hosted Clef · Needs an API key: a Cloudflare account ID in CLOUDFLARE_ACCOUNT_ID and a Cloudflare API token in CLOUDFLARE_AUTH_TOKEN

curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef -X POST -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" -d '{ "model": "clef", "state": "Checkout has been failing for every customer for the last hour.", "questions": { "urgent": { "type": "noul", "instructions": "Is this support request urgent?" }, "team": { "type": "choice", "instructions": "Which team should handle this request?", "criteria": { "billing": "Payments, invoices, and refunds", "technical": "Outages, errors, and configuration", "sales": "Plans and upgrades" } }, "severity": { "type": "score", "instructions": "How severe is the customer impact?", "criteria": ["No impact", "Minor", "Major", "Critical"] } } }'

Commands are copied from the official sources below. We have not run them.

Sources