I was bored.

Then I saw this announcement from Cloudflare. I didn't understand the hype behind Jev, Jeff, Clef… so I decided to have fun.

Decision models are…………

…classifiers. With a twist.

Under the hood it's a classifier: you send text plus a list of labels and get back a probability per label. What's new is that the labels arrive with the request (no retraining), one call can ask many questions at once, and the probabilities are calibrated, so you can threshold on them in code.

Compared with calling a chatbot and parsing its reply, there's no prompt-engineering for JSON and no off-list answers like "banana". It's one GPU pass instead of a generation loop, so about 60 ms instead of seconds.

"Decision model"Cloudflare's name, for Clef.
"System One model"TypeSafe's name, for Jev, after Kahneman's fast, intuitive System 1.
"Zero-shot classifier"What a researcher would call it: schema-conditioned, calibrated, multi-question.

They're the latest branch of a long family of classifiers:

  1. ~2018Fine-tuned classifier. BERT plus a fixed output head. Want a new label? Collect data and retrain.
  2. ~2019Zero-shot via NLI. Every label becomes a hypothesis ("this text is about billing"). One forward pass per label.
  3. ~2023LLM as classifier. Prompt, generate, parse. Flexible but slow, and it may answer "banana".
  4. ~2025Joint label encoders. GLiClass-style: the labels sit inside the input and are all scored in one pass.
  5. 2026Decision models. The same trick on a 9B LLM backbone: many typed questions at once, calibrated odds, images and long context.

The twist: the labels are input, not weights, so every request can bring its own. Every question in the form is scored jointly, in one pass. And training pushes the probabilities to be calibrated: of all the answers given at 0.9, about 90% should be right.

An LLM loops.It predicts a token, appends it and predicts again, trying to sound like the person who'd answer. Sometimes it thinks out loud first.A chat LLM runs in a loop: predict the next token, append it, repeat, until it has written something that reads like a human answer. for token in range(512): think(); mimic_you(); emit()
A decision model doesn't.One pass in, only the numbers out. One forward pass, then a probability per allowed answer. It fills in the bubbles. {"technical": 0.961, "billing": 0.039}

So I built a demo running on Scaleway with Clef-flash.

Here's what it is, and what happens to a request from the moment it arrives.

Meet Clef

Clef is an open-weights model (Apache-2.0, free to self-host) released by Cloudflare. Clef-flash is the 9B-parameter version: about 19 GB of weights, so it fits on one 48 GB GPU. That's the one running here, on a Scaleway L40S. It speaks the same HTTP API as TypeSafe's Jev.

Clef is Cloudflare's family of open decision models (Apache-2.0, on Hugging Face), and they speak TypeSafe's Jev API. There are two sizes: Clef (27B, Qwen3.8 backbone) and Clef-flash (9B, Qwen3.5 backbone), which is the one running here. Each one is an LLM trained to read well rather than write well, plus a small head that turns what it read into one score per allowed answer. Training rewards correct answers and punishes over-confidence (a Brier loss, then RLCD, Reinforcement Learning for Calibrated Decisions), so a 0.9 behaves like a 0.9.

9Bparameters (Clef-flash)
19 GBweights in BF16
64kcontext (we cap at 40k)
3question types: choice · score · noul
38.8 msmedian latency (Cloudflare's GPUs)
91.8MMLU (model card)

State can be text, JSON, images or video. choice picks a named option, score picks an ordered level (you also get the expected value) and noul is yes/no.

The stack, one request at a time

01HAProxy

Arrive

HAProxy is the bouncer at the door. It does the HTTPS handshake, turns away anyone sending too much (30 decisions a minute without a key) and refuses bodies over 20 MB. Keys get hashed before they're counted, like passwords, so the bouncer never holds the real thing.

Standard edge setup, no LLM magic yet: HAProxy terminates TLS, rate-limits per API key (or per IP when there's no key) and forwards to a single backend. That backend is one GPU box in Paris.

POST /v1/systemone
{ "model": "clef-flash",
  "state": "Checkout is down, orders blocked.",
  "questions": {
    "team":   { "type": "choice",
                "criteria": { "billing": "Payments",
                              "technical": "Bugs or outages" } },
    "urgent": { "type": "noul" } } }
02FastAPI

Admit

FastAPI is the receptionist. It checks your key against a list of hashes, then checks your form is filled in properly: real question types, not too many options, not too long. Jev clients ask for jev-latest; we nod and hand them Clef-flash.

Plain request validation: auth against hashed keys, JSON shape checks and size limits. Keyless requests get tighter limits. jev-latest is only a model-name alias, so existing Jev clients work as-is.

key  → sha256 → keys.json         ✓ "demo"
none → "public" (8k tokens, 16 questions, 1 image)
model "jev-latest" → clef-flash
03Tokenizer

Ingest

Think template rendering. Your state and every question with its allowed answers are written into one long string. That's the trick: the labels are part of the input, not baked into the model. While rendering, the server notes where each answer starts and ends, like keeping array indices. Images get chopped into small tiles that the model reads like words. If your text is too long it gets trimmed; the questions never are.

LLMs don't read JSON, they read tokens: integer IDs for chunks of text, roughly 4 characters each. So we render your state and every question, with its allowed answers, into one prompt string, tokenize it, and remember which token ranges belong to which answer. Images become tokens too: they're split into patches, and each patch enters the model like a word would.

<|im_start|>system
Read the complete state and schema. Decide every
field jointly. Each answer must be exactly one of
that field's allowed options.<|im_end|>
<|im_start|>user
STATE:
Checkout is down, orders blocked.

SCHEMA FIELDS:
FIELD 1
ID: team
TYPE: choice
INSTRUCTION: team
ALLOWED OPTIONS:
OPTION 1: {"description":"Payments","option_id":"billing"}
OPTION 2: {"description":"Bugs or outages","option_id":"technical"}
END FIELD
…
<|im_start|>assistant
<think>

</think>

JOINT SCHEMA DECISIONS:
04Batcher

Batch

A bus that waits 5 ms at the stop: whoever shows up in that window rides the same GPU trip, up to 16 passengers. Busy minutes cost fewer trips than requests.

GPUs are throughput machines: one call carrying 8 requests costs barely more than a call carrying 1. So requests wait up to 5 ms in a queue and get padded into a single tensor, up to 16 per call.

t=0.0ms  req A ┐
t=1.8ms  req B ├─ one batch → GPU
t=3.1ms  req C ┘
t=5.0ms  window closes
05L40S GPU

Infer

The backbone is a 9B-parameter function that turns every token into a list of 4,096 numbers describing what it means in context. It runs once over the whole text: no loop, nothing generated, nothing kept for later. Its speed trick: three out of four layers keep a running summary of what came before instead of re-reading everything ("linear attention"), and every fourth layer does a full re-read.

The head is a small add-on that sits the actual exam. It boils each question and each answer down to one vector, lets every answer search your text for evidence, lets the questions compare notes ("if it's an outage, severity is probably high"), then gives every answer a score. Picture a tiny jury working from a case file the backbone wrote.

A chat LLM runs its whole network once per output token, in a loop: a 200-word answer is about 300 sequential passes. Here we run it once over the input and stop. That's called "prefill only".

What comes out isn't text. It's a vector of 4,096 numbers for every input token (the model's reading of that token in context). A small extra network, the head, compares the vectors of each question with those of each allowed answer and outputs one score per answer. Think of it as a learned similarity function, trained so the right answer scores highest.

hidden = qwen(tokens)          # [batch, seq, 4096]
q, opts = pool(hidden, spans)  # one vector each
opts = route(opts, hidden) ×2  # find evidence
fields = joint(q, hidden)  ×4  # questions agree
logit  = prior + match(fields, opts)
06FastAPI

Respond

Scores become percentages that add up to 100% per question (that's the softmax). You get JSON back: the winner, every option's odds and the timing. A background thread writes the usage row, so your request never waits on the database.

A softmax turns each question's scores into probabilities that sum to 1. "Calibrated" means you can trust them as rates: of all the answers given at 0.9, about 90% are correct, so thresholds in your code mean something. Then it's a normal JSON response, with usage logged asynchronously.

{ "answers": {
    "team":   { "choice": "technical", "confidence": 0.961,
                "probabilities": { "billing": 0.039,
                                   "technical": 0.961 } },
    "urgent": { "noul": 0.93 } },
  "usage":  { "input_tokens": 141, "output_tokens": 0 },
  "timing": { "compute_ms": 61.4, "batch_size": 1 } }

Observations

~400-token decisionL4 (24 GB)L40S (48 GB)
GPU time167 ms61 ms
Max throughput6.3 req/s15.6 req/s
Hot dog photoI went to sleep84 ms
Busy web page (3k tokens)I went to sleep~300 ms

Cloudflare quotes a 38.8 ms median on its own GPUs, so we're in the same postcode.

  • GPU choice matters: 167 ms per request on an L4, 61 ms on an L40S.
  • The model needs an optional GPU-kernel package (plus gcc in the container) to be fast. Without it, it silently falls back to a slow path.
  • GPU kernels tune themselves the first time they see a new input size (about 10 s), so the server sends dummy requests at boot.
  • Conclusion: fast, and oddly good at boring decisions, which is most decisions.
  • Without the flash-linear-attention kernels, those 24 DeltaNet layers fall back to slow PyTorch.
  • Triton compiles C at runtime, so the container needs gcc. Surprise.
  • Kernels autotune the first time they see a new input size (10 s!), so the server warms up at boot and caches the results.
  • Conclusion: fast, and oddly good at boring decisions, which is most decisions.

And then I thought I'd show people.

Hot Dog Detector GPU
★★★★
🌭
[ Vision / Effect ]

When a photo is summoned: reveal the probability that it's a hot dog. Cannot be fooled by a bratwurst (probably).

ATK/84msDEF/510tok
XAV-EN001⌥1
Slack Exchange Trap 罠
[ TRAP CARD ⟲ ]
📅
[ Continuous Trap ]

Activate when an opponent plays a 60-minute invite with no agenda. Negate it and send it to the Slack thread.

SPOILER/should've been
XAV-EN002⌥2
Unclutter 魔
[ SPELL CARD ∞ ]
🧹
[ Continuous Spell ]

While face-up, every ad and cookie banner scored ≥ 0.9 is banished from your browser. Requires kitze's extension.

COST/1 API key
XAV-EN003⌥3
API Console DEV
★★★
⌨️
[ Developer / Normal ]

A plain JSON textarea that has seen things. Speaks fluent /v1/systemone and never generates a single token.

ATK/61msDEF/0 tokens out
XAV-EN004⌥4

Try it from your terminal

The demos need no key. For the API, ask me (@saalik) for an API key.
This runs on my pro account. Please don't fire me.