AI RACE— The AI Race
New Models

Cloudflare Unveils Clef Decision Models to Automate AI Agent Actions at Edge Speeds

Cloudflare has released Clef and Clef-flash, two open-source decision models built on Qwen that assign probabilities to predefined choices in as fast as 39 milliseconds.

10/03/2026, 01:19
Cloudflare ra mắt dòng mô hình ra quyết định Clef, tự động hóa tác vụ của AI agent với tốc độ mạng biên

Cloudflare has introduced Clef and Clef-flash, two specialized decision models engineered to streamline how AI agents evaluate data and execute workflows. Designed to assign probabilities to predefined multiple-choice options rather than generate open-ended text, the models aim to remove human bottlenecks from agentic pipelines by enabling downstream systems to trigger actions automatically.

The launch places Cloudflare into direct competition with TypeSafe AI's Jev model, which pioneered the concept in mid-September, as well as OpenAI's Decisions API based on GPT-6 Luna. To ease adoption, Cloudflare has made Clef’s API fully compatible with Jev.

Bridging the Gap Between Classifiers and LLMs

Decision models target a specific structural challenge in agentic architectures. While traditional large language models (LLMs) excel at reasoning and tool execution, their responses are prone to variation, hallucination, and high latency. Conversely, traditional machine-learning classifiers run rapidly but require dedicated retraining whenever a new category is introduced.

Clef bridges this divide by evaluating an input across multiple predefined questions in a single call, returning confidence probabilities for each candidate option. In a customer service scenario, for instance, Clef can evaluate an incoming inquiry, gauge its urgency, determine the appropriate team, and output clear probabilities. Downstream code can then automatically route the ticket, trigger an alert, or escalate the case to a human worker if confidence falls below a designated threshold.

According to Cloudflare, this degree of reliability means "a human does not necessarily need to be in the loop for agentic decisions anymore."

Architecture, Benchmarks, and Edge Speed

Both models in the family are built on open-weight Qwen foundations: the primary Clef model uses Qwen3.8-27B, while Clef-flash relies on Qwen3.5-9B. Cloudflare leaves the core base weights intact during training, instead using proprietary synthetic data to train specialized auxiliary components that extract calibrated probabilities directly from the model’s internal activations.

Training incorporates Cloudflare's own variant of Reinforcement Learning for Calibrated Decisions (RLCD)—the method originally used by TypeSafe AI. Clef also introduces multimodal support, accepting both text and images, and features a 64,000-token context window, double the capacity of Jev.

Speed across Cloudflare's distributed edge infrastructure serves as the main differentiator:

  • Clef-flash records a median response latency of approximately 39 milliseconds.
  • Clef achieves a median response latency of roughly 209 milliseconds.
  • By comparison, TypeSafe AI's Jev clocks in at just over 524 milliseconds.

In Cloudflare's internal evaluations across 43 benchmarks, Clef-flash achieved comparable accuracy to Jev at a fraction of the response time, while Clef led overall decision quality. On the API Bank benchmark, Clef-flash scored 93.11 and Clef achieved 91.93, compared to 88.19 for Jev. On PhishNChips, Clef posted 79.60 and Clef-flash scored 75.05, outperforming Jev’s 62.55. On the When2Call benchmark, Jev retained an advantage at 80.97 against Clef's 72.37 and Clef-flash's 65.58.

In internal testing by Cloudflare's threat intelligence team, Clef fetched, rendered, and classified a website—evaluating risk factors like phishing probabilities—in 2.2 seconds. The company's fastest general-purpose LLM took 4.7 seconds for the same task while returning fewer categories.

Enterprise Fine-Tuning and Open Release

Both Clef and Clef-flash have been released on Hugging Face under the Apache-2.0 open-source license and are accessible on Cloudflare’s Workers AI serverless platform.

Alongside the models, Cloudflare is rolling out a reinforcement learning service allowing organizations to tailor Clef to custom tasks. Initially delivered through forward-deployed engineers, the workflow will eventually expand into a self-service pipeline:

  1. Production traffic requests are recorded via Cloudflare's AI Gateway.
  2. The logged interactions are evaluated in containerized sandboxes.
  3. Custom-trained weights are trained and redeployed directly on Workers AI, utilizing infrastructure from Replicate, which Cloudflare acquired in late 2025.

Cloudflare plans to use the Clef models internally to triage customer support messages, evaluate automated abuse reports, and distinguish benign web crawlers from malicious automated bots.

◗ Sources

The Decoder10/03

Related stories