Cloudflare releases Clef and Clef-flash models for AI agent decision-making
Cloudflare has released the Clef and Clef-flash models, which assign probabilities to predefined options. They support both text and images and are available on the Workers AI platform as well as under the Apache-2.0 license on the Hugging Face platform.
Cloudflare has released the decision models Clef and Clef-flash. Instead of a long text response, they return a classification with probabilities for predefined options. Downstream code can, for example, evaluate the urgency of a customer request, pass it to the appropriate team, trigger an escalation, or involve a human. According to Cloudflare, this allows agents to make some decisions without ongoing human involvement.
According to Cloudflare, the Clef model is based on the Qwen3.8-27B model, and the Clef-flash model on the Qwen3.5-9B model. Both support text and images; the Clef model has a context window of 64 000 tokens. The company states a median latency of approximately 39 milliseconds for the Clef-flash model and 209 milliseconds for the Clef model, compared to more than 524 milliseconds for the Jev model. These are results of Cloudflare's own measurement. The API maintains full compatibility with the Jev model's API, which is intended to make the transition easier.
Both models run on the Workers AI platform and are available on the Hugging Face platform under the Apache-2.0 license. Cloudflare is simultaneously introducing a service for customizing the Clef model to specific tasks using reinforcement learning. Initially, specialized engineers are to help customers; a self-service platform is only planned.
Why it matters
Developers can directly connect the probabilities of individual options to further steps of the agent. For customer support, this offers automatic routing of requests, escalation of urgent cases, and handoff to a human. The compatible API may make it easier to try the models in applications using the Jev model; the stated speed advantages are so far supported only by Cloudflare's own measurement.
Release card
Clef
Cloudflare
- Context
- The context window is 64 000 tokens.
- Inputs
- The model accepts text and images and returns a classification with probabilities for predefined options.
- Licence
- Apache-2.0
- Availability
- The model is available on the Workers AI platform and on the Hugging Face platform.
- Medián latence Přibližně 209 milisekund Cloudflare states this value as the median response time of the Clef model in its own measurement.
- API Bank 91.93 The value expresses the accuracy of the Clef model on the API Bank benchmark according to Cloudflare's measurement.
- When2Call 72.37 The value expresses the accuracy of the Clef model on the When2Call benchmark according to Cloudflare's measurement.
- PhishNChips 79.60 The value expresses the accuracy of the Clef model on the PhishNChips benchmark according to Cloudflare's measurement.
- The model assesses the urgency of customer requests.
- The model determines the team to handle the request.
- The model classifies web pages.
- The model returns a decision from predefined options, so it does not serve for free-form text generation.
Cloudflare presents the Clef model as a direct alternative to the Jev model from TypeSafe AI. The source states a compatible API, lower median latency, image support, and double the context window compared to the Jev model.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
A developer can use the Clef and Clef-flash models in a project to obtain predefined decisions with probabilities and connect them to further steps of their agent.
For a business
Customer support can use the models to sort requests by urgency and responsible team, including escalation or handoff to a human.
ProcessesCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.