
Cloudflare has released Clef and Clef-flash, two open-weight models designed for the small but frequent decisions AI agents make between larger reasoning steps.
Instead of generating prose, the models evaluate application state and return typed answers with probabilities. An agent could use them to decide whether a support request is urgent, select the correct workflow, score an action against a safety policy, or determine when human approval is required.
Clef is a 27-billion-parameter model optimized for precision, while the 9-billion-parameter Clef-flash targets latency-sensitive workflows. Both offer a 65,536-token context window, support visual inputs, and use the Jev-compatible System One API. Cloudflare has released their weights under Apache 2.0 and made hosted versions available through Workers AI.
Across ten decision benchmarks selected by Cloudflare, either Clef or Clef-flash produced the highest score on seven. Cloudflare’s own tests also measured median decision latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, compared with 524.1 milliseconds for Jev. These are vendor-run results and have not yet received broad independent replication.
Hosted pricing is $0.24 per million input tokens for Clef and $0.09 per million for Clef-flash. Cloudflare is also introducing reinforcement-learning fine-tuning, initially through its engineering team, with a self-service platform planned later.
Why it matters
Most production agents do not need a frontier language model for every routing, classification, or guardrail decision. A smaller model that returns bounded probabilities can make those steps faster, cheaper, and easier to validate.
That creates a practical two-layer architecture: use Clef for rapid decisions in the agent’s execution path, then call a general-purpose model only when deeper reasoning or content generation is needed. The pattern can complement independent policy enforcement and sandbox controls outside the agent, separating fast model-based judgment from hard runtime boundaries.
The models’ open weights also let teams test this design locally before adopting Cloudflare’s hosted infrastructure.