
What happened
A developer testing Cloudflare's Clef and Clef-flash, released October 1 under Apache 2.0, found Jev's official Python SDK ran on Clef after only changing the endpoint and model name, with identical answer structure.
Why it matters
The two families of decision models look interchangeable on the code level, so teams already using Jev can compare them on their own data rather than rewriting an application to try Clef.
What to watch
Whether the speed case holds depends on region and network, since the official 38.8 ms figure is a Cloudflare-internal median; the test also found Jev stays the cheapest option per judgment.
WHO IT HITSTeams building AI agents or support-triage tools that call a small decision step on every request, and platform or ML engineers weighing Cloudflare against TypeSafe, now have a same-code swap path but must benchmark speed and cost on their own network.
Summaries like this, in your inbox every morning.
Clef is the first model Cloudflare's Workers AI team trained itself, built on Qwen3.8-27B and Qwen3.5-9B backbones and released alongside a reinforcement-learning fine-tuning pipeline. That pipeline stitches together AI Gateway, Workers AI, Containers and a new Trainer product, but only a design-partner program with Cloudflare's field engineers is open today, with no self-service timing or pricing given.
The benchmark picture is mixed rather than a clean win. On the community Jev Decision Index, Cloudflare reports Clef in first place, but the leaderboard carries a "Self-reported" note and the scores have not yet been reproduced on the original community board. Clef leads on classification and intent tasks such as BANKING77 and API-Bank, while Jev is far ahead on knowledge and multi-step reasoning such as GPQA Diamond and BBH. Clef-flash is fast but loses considerable accuracy on out-of-scope detection and hallucination detection, so it is not simply a faster Clef.
The clearest differentiators appear to be image input, which Jev does not support, and the openly licensed weights that let a team run both models locally with the same answers as the cloud. What the test ultimately suggests is that the gap between these models is smaller than the gap made by how well instructions and criteria are written, so the choice is likely to hinge less on the model and more on whether a team has already documented its own decision rules and threshold behavior.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
NetApp and Iterate.ai are packaging the AIPod Mini with Iterate's Generate platform and an embedded LLM, so en…
Mercor had 12 licensed CPAs work through simplified APEX Accounting Benchmark tasks

Testing Azure API Management's llm-token-limit policy at 800 tokens per hour, actual consumption hit 1,472 tok…

Qwen released Qwen3.8-Flash-Next on August 27, 2026, calling it a preview of the architecture planned for Qwen…

A Zenn article floated a hackathon where participants get the theme on the day, use no PC, internet, smartphon…

Anthropic's Message Batches API offers a 50% off rate, takes up to 10,000 requests per batch, and returns resu…
