Cloudflare Introduces Clef and Clef-flash Open-Weight Decision Models
Cloudflare's Workers AI team has released Clef and Clef-flash, open-weight decision models that output typed probabilities instead of free-form text.

Cloudflare has unveiled Clef and Clef-flash, marking the first time its Workers AI team has trained and released proprietary models. Rather than operating as standard chatbots that generate text sequentially, these models function as specialized decision engines. Users provide an input state alongside a schema of typed questions, and the system responds directly with calculated probabilities for every permitted answer, removing the need for manual text parsing.
Both models are available as open-weight releases under the Apache 2.0 license and maintain compatibility with TypeSafe AI’s Jev API. Developers can deploy them immediately on Workers AI, or access the weights on Hugging Face for independent self-hosting.
Architecture and Mechanics
The fundamental design of a decision model shifts focus away from token generation. Instead of writing out responses, Clef evaluates a fixed collection of inquiries against provided data. The framework supports three distinct question formats: binary yes/no parameters, named choice selections with specific confidence metrics, and ordinal scores evaluated against an ordered rubric. On Workers AI, a single execution call can handle up to 64 individual questions and accept up to 4 images.
Clef is built upon the Qwen3.8-27B backbone, while Clef-flash derives from the Qwen3.5-9B model. Both versions retain their underlying vision encoders. The inference process executes in two phases. First, the main backbone processes the state and question inputs in a single prefill pass. Next, a lightweight joint schema head—implemented as a small transformer—reads the resulting hidden states, routes relevant evidence to each question, allows fields to cross-attend, and scores all possible options simultaneously. A per-question softmax operation subsequently converts the raw logits into final probabilities.
During training, the primary backbones remained frozen while the routing head was jointly optimized using rank-256 low-rank adapters. The optimization loss function pairs label-smoothed cross-entropy alongside a Brier loss to ensure calibration. Additionally, an objective called Reinforcement Learning for Calibrated Decisions awards partial credit when adjacent ordinal choices are selected.
Performance and Benchmarks
Evaluation data provided by Cloudflare across ten benchmarks from the Decision Index 0.2.1 suite shows Clef outperforming competitors in seven categories. For instance, on BANKING77 macro-F1 tests, Clef achieved 94.20 compared to Jev's 79.74, and on home appliances case-exact matching, Clef-flash reached 97.73 against Jev's 52.27. In a threat intelligence trial, Clef successfully classified a domain in 2.2 seconds, whereas gpt-oss-120b required 4.7 seconds. Median latency tests registered 209.3 ms for Clef and 38.8 ms for Clef-flash.
However, Jev maintains advantages in knowledge-intensive examinations such as GPQA Diamond, MMLU-Pro, and BBH. Cloudflare notes that all performance figures are currently vendor-reported and await independent verification.
What it means for developers
For engineers building automated pipelines, structured data validation, or classification systems, decision models offer an alternative to traditional language models that require complex prompt engineering and post-processing parsers. Because these models output typed probabilities rather than strings, integration into existing software architectures becomes more deterministic.
Developers can easily test various top AI models cheaply through one API at https://apixoai.online. Furthermore, teams looking to deploy Clef or Clef-flash can access them via the Workers AI binding, REST API, or AI Gateway, with self-hosting options tested on single H200 hardware setups.
Source: Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text — MarkTechPost. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

