I was bored.
Then I saw this announcement from Cloudflare. I didn't understand the hype behind Jev, Jeff, Clef… so I decided to have fun.
Decision models are…………
…classifiers. With a twist.
Under the hood it's a classifier: you send text plus a list of labels and get back a probability per label. What's new is that the labels arrive with the request (no retraining), one call can ask many questions at once, and the probabilities are calibrated, so you can threshold on them in code.
Compared with calling a chatbot and parsing its reply, there's no prompt-engineering for JSON and no off-list answers like "banana". It's one GPU pass instead of a generation loop, so about 60 ms instead of seconds.
They're the latest branch of a long family of classifiers:
- ~2018Fine-tuned classifier. BERT plus a fixed output head. Want a new label? Collect data and retrain.
- ~2019Zero-shot via NLI. Every label becomes a hypothesis ("this text is about billing"). One forward pass per label.
- ~2023LLM as classifier. Prompt, generate, parse. Flexible but slow, and it may answer "banana".
- ~2025Joint label encoders. GLiClass-style: the labels sit inside the input and are all scored in one pass.
- 2026Decision models. The same trick on a 9B LLM backbone: many typed questions at once, calibrated odds, images and long context.
The twist: the labels are input, not weights, so every request can bring its own. Every question in the form is scored jointly, in one pass. And training pushes the probabilities to be calibrated: of all the answers given at 0.9, about 90% should be right.
for token in range(512):
think(); mimic_you(); emit(){"technical": 0.961,
"billing": 0.039}So I built a demo running on Scaleway with Clef-flash.
Here's what it is, and what happens to a request from the moment it arrives.
Meet Clef
Clef is an open-weights model (Apache-2.0, free to self-host) released by Cloudflare. Clef-flash is the 9B-parameter version: about 19 GB of weights, so it fits on one 48 GB GPU. That's the one running here, on a Scaleway L40S. It speaks the same HTTP API as TypeSafe's Jev.
Clef is Cloudflare's family of open decision models (Apache-2.0, on Hugging Face), and they speak TypeSafe's Jev API. There are two sizes: Clef (27B, Qwen3.8 backbone) and Clef-flash (9B, Qwen3.5 backbone), which is the one running here. Each one is an LLM trained to read well rather than write well, plus a small head that turns what it read into one score per allowed answer. Training rewards correct answers and punishes over-confidence (a Brier loss, then RLCD, Reinforcement Learning for Calibrated Decisions), so a 0.9 behaves like a 0.9.
State can be text, JSON, images or video. choice picks a named option, score picks an
ordered level (you also get the expected value) and noul is yes/no.
The stack, one request at a time
Arrive
HAProxy is the bouncer at the door. It does the HTTPS handshake, turns away anyone sending too much (30 decisions a minute without a key) and refuses bodies over 20 MB. Keys get hashed before they're counted, like passwords, so the bouncer never holds the real thing.
Standard edge setup, no LLM magic yet: HAProxy terminates TLS, rate-limits per API key (or per IP when there's no key) and forwards to a single backend. That backend is one GPU box in Paris.
POST /v1/systemone
{ "model": "clef-flash",
"state": "Checkout is down, orders blocked.",
"questions": {
"team": { "type": "choice",
"criteria": { "billing": "Payments",
"technical": "Bugs or outages" } },
"urgent": { "type": "noul" } } }
Admit
FastAPI is the receptionist. It checks your key against a list of hashes, then checks your form is filled in properly: real question types, not too many options, not too long. Jev clients ask for jev-latest; we nod and hand them Clef-flash.
Plain request validation: auth against hashed keys, JSON shape checks and size limits. Keyless requests get tighter limits. jev-latest is only a model-name alias, so existing Jev clients work as-is.
key → sha256 → keys.json ✓ "demo"
none → "public" (8k tokens, 16 questions, 1 image)
model "jev-latest" → clef-flash
Ingest
Think template rendering. Your state and every question with its allowed answers are written into one long string. That's the trick: the labels are part of the input, not baked into the model. While rendering, the server notes where each answer starts and ends, like keeping array indices. Images get chopped into small tiles that the model reads like words. If your text is too long it gets trimmed; the questions never are.
LLMs don't read JSON, they read tokens: integer IDs for chunks of text, roughly 4 characters each. So we render your state and every question, with its allowed answers, into one prompt string, tokenize it, and remember which token ranges belong to which answer. Images become tokens too: they're split into patches, and each patch enters the model like a word would.
<|im_start|>system
Read the complete state and schema. Decide every
field jointly. Each answer must be exactly one of
that field's allowed options.<|im_end|>
<|im_start|>user
STATE:
Checkout is down, orders blocked.
SCHEMA FIELDS:
FIELD 1
ID: team
TYPE: choice
INSTRUCTION: team
ALLOWED OPTIONS:
OPTION 1: {"description":"Payments","option_id":"billing"}
OPTION 2: {"description":"Bugs or outages","option_id":"technical"}
END FIELD
…
<|im_start|>assistant
<think>
</think>
JOINT SCHEMA DECISIONS:
Batch
A bus that waits 5 ms at the stop: whoever shows up in that window rides the same GPU trip, up to 16 passengers. Busy minutes cost fewer trips than requests.
GPUs are throughput machines: one call carrying 8 requests costs barely more than a call carrying 1. So requests wait up to 5 ms in a queue and get padded into a single tensor, up to 16 per call.
t=0.0ms req A ┐
t=1.8ms req B ├─ one batch → GPU
t=3.1ms req C ┘
t=5.0ms window closes
Infer
The backbone is a 9B-parameter function that turns every token into a list of 4,096 numbers describing what it means in context. It runs once over the whole text: no loop, nothing generated, nothing kept for later. Its speed trick: three out of four layers keep a running summary of what came before instead of re-reading everything ("linear attention"), and every fourth layer does a full re-read.
The head is a small add-on that sits the actual exam. It boils each question and each answer down to one vector, lets every answer search your text for evidence, lets the questions compare notes ("if it's an outage, severity is probably high"), then gives every answer a score. Picture a tiny jury working from a case file the backbone wrote.
A chat LLM runs its whole network once per output token, in a loop: a 200-word answer is about 300 sequential passes. Here we run it once over the input and stop. That's called "prefill only".
What comes out isn't text. It's a vector of 4,096 numbers for every input token (the model's reading of that token in context). A small extra network, the head, compares the vectors of each question with those of each allowed answer and outputs one score per answer. Think of it as a learned similarity function, trained so the right answer scores highest.
hidden = qwen(tokens) # [batch, seq, 4096]
q, opts = pool(hidden, spans) # one vector each
opts = route(opts, hidden) ×2 # find evidence
fields = joint(q, hidden) ×4 # questions agree
logit = prior + match(fields, opts)
Respond
Scores become percentages that add up to 100% per question (that's the softmax). You get JSON back: the winner, every option's odds and the timing. A background thread writes the usage row, so your request never waits on the database.
A softmax turns each question's scores into probabilities that sum to 1. "Calibrated" means you can trust them as rates: of all the answers given at 0.9, about 90% are correct, so thresholds in your code mean something. Then it's a normal JSON response, with usage logged asynchronously.
{ "answers": {
"team": { "choice": "technical", "confidence": 0.961,
"probabilities": { "billing": 0.039,
"technical": 0.961 } },
"urgent": { "noul": 0.93 } },
"usage": { "input_tokens": 141, "output_tokens": 0 },
"timing": { "compute_ms": 61.4, "batch_size": 1 } }
Observations
| ~400-token decision | L4 (24 GB) | L40S (48 GB) |
|---|---|---|
| GPU time | 167 ms | 61 ms |
| Max throughput | 6.3 req/s | 15.6 req/s |
| Hot dog photo | I went to sleep | 84 ms |
| Busy web page (3k tokens) | I went to sleep | ~300 ms |
Cloudflare quotes a 38.8 ms median on its own GPUs, so we're in the same postcode.
- GPU choice matters: 167 ms per request on an L4, 61 ms on an L40S.
- The model needs an optional GPU-kernel package (plus
gccin the container) to be fast. Without it, it silently falls back to a slow path. - GPU kernels tune themselves the first time they see a new input size (about 10 s), so the server sends dummy requests at boot.
- Conclusion: fast, and oddly good at boring decisions, which is most decisions.
- Without the
flash-linear-attentionkernels, those 24 DeltaNet layers fall back to slow PyTorch. - Triton compiles C at runtime, so the container needs
gcc. Surprise. - Kernels autotune the first time they see a new input size (10 s!), so the server warms up at boot and caches the results.
- Conclusion: fast, and oddly good at boring decisions, which is most decisions.
And then I thought I'd show people.
When a photo is summoned: reveal the probability that it's a hot dog. Cannot be fooled by a bratwurst (probably).
Activate when an opponent plays a 60-minute invite with no agenda. Negate it and send it to the Slack thread.
While face-up, every ad and cookie banner scored ≥ 0.9 is banished from your browser. Requires kitze's extension.
A plain JSON textarea that has seen things. Speaks fluent /v1/systemone and never generates a single token.
Try it from your terminal
The demos need no key. For the API, ask me (@saalik) for an API key.
This runs on my pro account. Please don't fire me.