DEV Community

Cover image for Same Judge, Two Price Tags: Benchmarking Jev Against Cloudflare Open-Source Clef
HIROKI II
HIROKI II

Posted on

Same Judge, Two Price Tags: Benchmarking Jev Against Cloudflare Open-Source Clef

You ask an assistant to sort 30 customer messages into three categories. Every reply comes back as a five-paragraph essay with the verdict buried in the last line. You copy conclusions into a spreadsheet by hand, re-checking each one.

That is not a prompting failure. The assistant is treating a judgment task like a writing task. A different class of model — decision models — exists for this: pick one option from a fixed list, attach a confidence score, write nothing else. TypeSafe Jev pioneered the category. On October 1, 2026, Cloudflare entered it with two open-source models, Clef and Clef-flash, on Workers AI, Jev-API compatible, Apache 2.0 on Hugging Face.

All figures below come from official pages verified October 5, 2026. I have not called either API.

Try the concept in any chat window

Take ten real judgment tasks. Window one: judge and explain. Window two: constrain it —

Categories: refund, exchange, inquiry.
Pick exactly one. Confidence 0-100.
No answers outside the list.
Enter fullscreen mode Exit fullscreen mode

Window two output drops straight into a spreadsheet. That is the whole idea of a decision model: the model chooses; your code acts.

The bill

Item Jev 1.13 Clef Clef-flash
Median latency 524.1 ms 209.3 ms 38.8 ms
p95 536.0 ms 238.6 ms 122.4 ms
$/M input tokens $0.042 $0.240 $0.090
Context 64k per request (state + longest question 32k) 64k 64k
Image input No Yes Yes
Open weights No Apache 2.0 Apache 2.0

Three observations the marketing copy blurs.

2.5x-13.5x faster is two models against one. Clef regular is 2.5x; Clef-flash is 13.5x. Cite them separately.

Open is not cheap. On Workers AI, Clef costs 5.7x Jev per million input tokens; Clef-flash still costs 2.1x. The free daily 10,000 Neurons cover roughly 460k Clef input tokens, so light experiments cost nothing — sustained volume does not favor Clef. The Apache 2.0 license buys you the third option: self-host the weights, where the token meter disappears and the GPU bill begins. Backbones are Qwen3.8-27B (Clef) and Qwen3.5-9B (Clef-flash).

Small scale is noise. 100k messages x 500 input tokens: Jev about $2.10, Clef about $12.00, Clef-flash about $4.50.

On quality: Cloudflare blog reports Clef ahead on five public suites (BFCL 98.47 vs 95.75, API-Bank 91.93 vs 88.19, ToolRet 69.19 vs 65.28), but on TypeSafe own four workflows Clef wins only three of four. Every number is vendor-reported; no independent replication exists yet.

Three deployment tiers

  • Jev hosted: POST /v1/systemone; managed limits (100K tokens/s, 80 req/s).
  • Clef on Workers AI: one REST call with state and questions; switch from Jev by configuration.
  • Clef self-hosted: Hugging Face weights; you own the GPU.

Skip for now

  • Don't migrate a live pipeline before finishing the two-window exercise.
  • Don't self-host 27B on a laptop; start with Clef-flash on a real GPU box.
  • Don't treat vendor benchmarks as promises.
  • Skip the RL fine-tuning product until you have volume.

Sources: Cloudflare blog, Workers AI pricing, TypeSafe models, HF: Cloudflare/clef.

Top comments (0)