DEV Community

Shouvik Palit
Shouvik Palit

Posted on

Clef-Flash: the 9B model that decides instead of chats (I spotted it on DEV·TV)

I was half-watching DEV·TV, the little retro TV I built that plays dev news on autopilot, when the Hugging Face channel landed on Cloudflare/clef-flash: 275 likes and 1,303 downloads a couple of days after release. I'd never heard of it, so I stopped the channel and read what Cloudflare had published.

Here's what I found. I haven't run the model myself, so everything below comes from Cloudflare's announcement, the Hugging Face cards and the llama.cpp discussions, and I say whose claim each number is.

How it showed up

DEV·TV's Hugging Face channel pulls trending models created in the last 14 days and shows their likes and downloads. That's exactly the window where a brand-new model is easy to miss in a normal feed. Clicking the story opens the model card inside the TV, so I could read the card without leaving the picture.

What a decision model is

A chat model generates text token by token. A decision model skips the text.

You give it a state (text, JSON or an image) and a set of typed questions: yes/no, pick one option, or a score. It returns a probability for every allowed answer in a single forward pass. There's no free-form output to parse, no JSON wrapped in a code fence, and no retry loop.

Cloudflare's example is support triage: pass in a customer message and ask whether it's urgent and which team owns it. Your code routes the ticket from the probabilities, and sends low-confidence cases to a human.

Typesafe AI's Jev models started this category a couple of weeks earlier. Clef is Cloudflare's entry, and it's Jev/SystemOne API compatible, so it's meant as a drop-in swap.

The release

  • Clef-Flash: 9B, post-trained from Qwen3.5-9B, tuned for latency
  • Clef: the 27B sibling, tuned for precision
  • License: Apache 2.0, weights on Hugging Face, also hosted on Workers AI
  • Context: 64k tokens (Cloudflare contrasts this with Jev's 32k)
  • Vision: Cloudflare says Clef has a vision encoder, so it can classify images

The numbers (Cloudflare's own)

These come from Cloudflare's benchmark, not independent tests. Latency in milliseconds:

Model Median p95
Clef-Flash 38.8 122.4
Clef 209.3 238.6
Jev 524.1 536.0
Kev-9B 51.4 187.9
DiffusionGemma Jev 84.4 211.2
Laya 5.8 222.5

Cloudflare says its models beat the other decision models on latency except Laya, which it says trades quality for speed.

Quality is more mixed for Flash. It does well on some tests (BANKING77 macro-F1 of 90.9 vs Jev's 79.7) and worse on others (When2Call: 65.6 vs Jev's 81.0, and CLINC150+OOS: 66.8 vs 89.3). Pick the model against your task, not the headline.

One caveat: those latencies come from Cloudflare's hosted setup. A third-party tool I found listed Clef-Flash at 532 ms for five questions on an RTX 4090, but that's a different method, so don't compare the two directly.

Running it locally

It isn't a normal Ollama chat model. I couldn't find evidence that stock Ollama serves the decision endpoint, and a community MLX card warns that generic chat runners load the backbone but produce meaningless text.

The route that exists today is llama.cpp's llama-server, which has a /v1/systemone endpoint:

  • Native Clef support was a draft PR (#29831) on Oct 1. A GitHub issue I saw says text support has since merged, text only. [verify before publishing]
  • GGUF quants are available. A Q4_K_M is about 5.8 GB in one community quant.
  • Image input isn't supported in llama.cpp yet; it's tracked separately.
  • Apple Silicon users have a community MLX conversion.

A request would look roughly like this. I haven't tested it, so check the model card for the exact shape:

llama-server -hf ggml-org/Clef-Flash-GGUF:Q4_K_M

curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "clef-flash",
    "state": "Disk on db-02 is at 97% and growing about 1% an hour.",
    "questions": {
      "page_oncall": { "type": "noul", "instructions": "Should someone be paged right now?" },
      "severity": {
        "type": "score",
        "instructions": "How severe is this?",
        "criteria": ["Info", "Minor", "Major", "Critical"]
      }
    }
  }'
Enter fullscreen mode Exit fullscreen mode

Where it would fit

Good uses are hot-path decisions that need a fast, typed answer: ticket routing, intent detection, a gate in front of an expensive LLM call, or alert triage. It's the wrong tool for open-ended writing or reasoning.

An idea I haven't built: a model like this could tag or rank DEV·TV's stories, for example "is this relevant to web developers?" or "is this a security story?", without a chat model's cost or its parsing headaches. That's speculation, not a feature.

What I'm watching

  • Whether Ollama adds support, or a decision-model runtime becomes the standard way to serve these
  • Independent benchmarks of Clef-Flash against Jev on real workloads
  • Image support in llama.cpp

If you've run Clef locally, I'd like to hear your latency and hardware.


I spotted this on DEV·TV, my open-source dev-news TV. It's static files with no backend and no API keys, and it pulls GitHub, Hacker News, Hugging Face, papers, CVEs and more into 12 channels. If it's useful to you, a ⭐ on GitHub helps a lot: github.com/shouvik12/devtv

Top comments (0)