Three companies shipped the same idea in about 48 hours. On October 1, Cloudflare released Clef and Clef-flash, AWS published Strands Decider 2B, and Perplexity opened a Decisions API on its own open 27B model. None of them write prose. They read some input, answer a fixed question and return a probability.
That sounds small, but it changes the math on AI for decision-making inside real operations. Most automated workflows don't need an essay. They need "is this a refund request?", "which queue?", "does this order look like fraud?", answered fast, cheaply and in a format your code can trust.
TL;DR#
- A new class of "decision models" returns typed answers (yes/no, pick one, score) with calibrated probabilities instead of generated text.
- Vendors claim latency in the tens to hundreds of milliseconds and input prices around $0.04 to $0.24 per million tokens, with free output. Benchmarks are mostly vendor-reported.
- If your workflows already use a chat LLM for classification or routing, it's worth piloting one of these on a single high-volume decision this quarter.
What actually shipped this week#
The category was kicked off by TypeSafe, which launched Jev on September 15 and called it a "System One" model, borrowing Kahneman's term for fast, intuitive thinking. Two weeks later, the big platforms answered.
Cloudflare Clef and Clef-flash#
Clef is a 27B model and Clef-flash a 9B model, both built on Qwen backbones and released as open weights under Apache 2.0. They run on Workers AI, answer yes/no, multiple-choice and ordinal scoring questions, and accept text, JSON, images and video. Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, with pricing of $0.24 and $0.09 per million tokens. It also announced a reinforcement learning fine-tuning service, starting with hands-on engagements before going self-serve.
AWS Strands Decider 2B#
AWS took a 2B base model, removed the text-generation head and attached a tiny pointer head that compares the model's internal state against each option in a single pass. It reports a 115 ms median on a consumer RTX 3090 and says it runs on CPU and Apple silicon too. Weights, code and the training recipe are Apache-2.0, installable with pip install strands-decider.
Perplexity Decisions API#
Perplexity's API runs pplx-decider-v1-27b, takes up to 128 questions per request and charges $0.04 per million input tokens, with free output. The docs say short inputs come back in under two seconds.
Why decision models matter for AI workflows#
Most AI workflows we see in growing businesses look like this: a chat model gets a ticket or an order, is told "reply with one of these five labels", and a parser hopes the answer comes back clean. It usually does. The failures are the expensive part: extra words, a made-up category, a timeout on a busy day.
Decision models remove that whole failure class. The answer is always one of the options you defined, because the model can't produce anything else. You also get a probability, which is the part people underrate. A probability lets you set a threshold: auto-route anything above 0.9, send the rest to a person. That's how you get automation you can actually defend to an operations manager.
Then there's cost. With output tokens free and input priced in cents per million, classifying every inbound message stops being a budget conversation. One Hacker News estimate quoted by Traictory put a million 300-token decisions at about $72 on Clef and $12.60 on Jev. Either number is small next to an hour of manual triage a day.
Technical breakdown and trade-offs#
Under the hood, the trick is the same across vendors: skip token-by-token generation and score the allowed answers directly. That's why latency drops and why output is "free".
A typical request looks like this (shape based on Perplexity's docs, simplified):
{
"model": "pplx-decider-v1-27b",
"state": { "ticket": "Box arrived crushed, want a replacement not a refund" },
"questions": {
"intent": { "type": "choice", "criteria": ["refund", "replacement", "tracking", "other"] },
"urgent": { "type": "noul", "instructions": "Customer is threatening a chargeback" }
}
}
You get back the chosen option, a probability per option and a yes/no probability. No parsing, no retries.
| Option | Size | Hosting | Reported median latency | Input price per 1M tokens |
|---|---|---|---|---|
| TypeSafe Jev | Not disclosed | Hosted only | 70 to 500 ms (vendor range) | $0.042 |
| Cloudflare Clef | 27B | Workers AI or self-host | 209.3 ms | $0.24 |
| Cloudflare Clef-flash | 9B | Workers AI or self-host | 38.8 ms | $0.09 |
| Perplexity decider | 27B | API or open weights | Under 2 s for short inputs | $0.04 |
| AWS Strands Decider | 2B | Self-host (CPU, GPU, Mac) | 115 ms on RTX 3090 | Your own compute |
The trade-offs nobody should skip#
- No explanations. You get a probability, not a reason. For regulated decisions (credit, claims, hiring) that's a real gap, and DataCamp's review flags it too.
- Valid isn't the same as correct. The answer always fits your schema, but it can still be the wrong option. Calibration and thresholds matter more than ever.
- Benchmarks are early. Cloudflare's own table shows Clef winning on intent classification (BANKING77, 94.20 vs 79.74) but losing badly on knowledge-heavy questions (GPQA Diamond, 48.0 vs 78.3). Treat every number here as vendor-reported until you test on your data.
- Answers must be known in advance. If the job is open-ended, you still need an LLM.
AI for decision-making: what it means for your business#
If your team handles high-volume, repetitive judgment calls, this is the cheapest, most predictable way yet to automate the "sort and route" layer. Good candidates:
- Support triage: intent, urgency, language and sentiment on every ticket, before a human sees it.
- Order and fulfilment checks: flag address mismatches, risky orders or items needing manual packing.
- Back-office data reconciliation: "does this invoice line match this PO line?" as a yes/no with confidence.
- Agent guardrails: a fast check before an AI agent sends an email or issues a refund.
- Model routing: decide whether a request needs an expensive LLM at all.
The self-hostable options help with privacy too: a 2B model on your own hardware keeps customer data in your environment.
How to act on it: a practical checklist#
- List your decisions. Write down every place a person or an LLM picks from a fixed set of answers. Rank by volume times cost of a mistake.
- Pick one. Start with a decision that's high-volume and low-risk, like ticket intent.
- Build a labelled test set. Pull 300 to 500 real historical examples with the correct answer. This matters more than the model choice.
- Run two or three models side by side. Measure accuracy, latency and cost on your set, not on public benchmarks.
- Set thresholds. Decide the confidence level for auto-action and route everything else to a human queue.
- Log every decision. Store input, answer, probability and outcome so you can audit and retrain later.
- Keep the LLM for the long tail. Use the decision model as the first pass and escalate the uncertain cases.
inbound ticket ──> decision model ──┬─ p > 0.9 ──> auto-route / auto-tag
├─ 0.6–0.9 ──> suggested label, human confirms
└─ p < 0.6 ──> LLM or human review
How MagicMakers Lab approaches this#
This is exactly the layer our AI Integration & Automation work focuses on: putting small, reliable decisions inside the tools you already run, like Zoho, Shopify or Stripe, with confidence thresholds, human approval for edge cases and a log of every call. It's the same thinking behind the Framico build, where 200+ orders a day ship with zero manual steps and support went from five people to one. We'd pick the model last, after we've mapped your decisions and built the test set.
Key takeaways#
- Decision models return a fixed answer plus a probability, so the output can't drift off-schema.
- Cloudflare, AWS and Perplexity all shipped one within days of TypeSafe's Jev, and three of the four are open weights.
- The big wins are cost, latency and predictable integration, not raw intelligence.
- The gaps are missing explanations and benchmarks that are still mostly vendor-reported.
- Start with one high-volume decision, a real test set and a confidence threshold.
FAQ#
What is a decision model in AI?#
A decision model is an AI model that answers a fixed question instead of writing text. You give it an input and a set of allowed answers (yes/no, one of several options or a score), and it returns the chosen answer with a probability. Because it never generates free text, the output always fits your schema and is easy to plug into code.
How is AI used for decision-making in business?#
Most practical uses are operational: routing support tickets, flagging risky orders, checking documents, matching invoices to purchase orders and deciding when a human should step in. The safest setup lets AI handle confident, low-risk calls automatically, sends uncertain ones to a person, and logs every decision so the team can review accuracy over time.
Are decision models better than ChatGPT-style LLMs for classification?#
For fixed-label classification they're usually faster, cheaper and easier to integrate, because there's nothing to parse. They aren't smarter, though. Cloudflare's own results show its Clef model trailing on knowledge-heavy questions. If your task needs reasoning, explanations or open-ended answers, a general LLM is still the right tool. Test both on your own data.
Can I run an AI decision model on my own servers?#
Yes. Cloudflare's Clef models and AWS's Strands Decider 2B are released under Apache 2.0, and AWS says its 2B model runs on CPU, consumer GPUs and Apple silicon. Self-hosting keeps customer data in your environment, which helps with privacy rules, but you take on hosting, monitoring and updates yourself.
If your team spends hours a day sorting, tagging or checking things that follow clear rules, there's probably a decision in there a model can make for a fraction of a cent. We'll help you find it and tell you straight whether it's worth automating. Book a free audit