Three weeks ago "decision model" was not a product category. On September 15, 2026, TypeSafe AI launched Jev, a model that returns a typed answer with a probability instead of writing text. Since then Perplexity, OpenAI, Cloudflare, Liquid AI, the Strands team, Superagent, Fastino and a pile of independent developers have shipped their own. This is a ranked list of ten you can use now, with the price, the speed and the one caveat that matters for each.
If you want the concept first, read our practitioner guide to decision models. The one-sentence version: a decision model reads your input and a closed set of answers, and returns a probability for each in a single forward pass, so you pay for input only and never parse a malformed response.
TL;DR: the ranked list
| Rank | Model | Maker | Access | Headline number | Best for |
|---|---|---|---|---|---|
| 1 | Jev | TypeSafe AI | Hosted, closed | $0.042 per million input tokens | The reference choice for text decisions |
| 2 | pplx-decider v1.1 | Perplexity | Open weights plus API | Decision Index 0.3: 61.56 vs Jev 57.9 | Strongest open option, multimodal |
| 3 | Decisions API (gpt-6-luna) | OpenAI | Hosted, public beta | $0.10 per million input tokens | One-vendor shops and compliance |
| 4 | Clef and Clef-flash | Cloudflare | Open weights plus Workers AI | $0.24 and $0.09 per million input | Teams already on Cloudflare |
| 5 | d1, d1-3B, d1-omni-600M | Liquid AI | Hosted plus open weights | 16 ms on a Jetson AGX Thor (d1-3B) | Edge devices, images and audio |
| 6 | Strands Decider 2B | Strands team | Open weights | About 115 ms on an RTX 3090 | Retrainable local agent hops |
| 7 | Security-One 27B | Superagent | Open weights plus API | 599 of 600 BIPIA attacks caught | Prompt-injection screening |
| 8 | GLiNER2.5-Decide | Fastino | Open weights plus API | 167 ms p50 on CPU | Rule-bound decisions without a GPU |
| 9 | Julia 1 | Supersonic Labs | Open weights | 144M parameters, 550 MiB | Laptop and browser |
| 10 | Kev and Jeff | Community | Open weights | 0.8B to 9B fine-tunes | Training on your own categories |
The order is explainx.ai's judgment for a builder picking something this month, weighing maturity, openness, documentation and how much independent testing exists. It is not a benchmark ranking. Every accuracy figure below is a vendor or author claim that we did not re-run, and the vendors use different benchmarks that are not comparable with each other.
How we ranked them
Four questions decided the order.
- Can I use it today? A model you can call or download beats an announcement.
- Is there independent evidence? Hacker News tests, third-party plugins and corrections count. A vendor panel alone does not.
- Can I leave? Open weights and a shared API shape reduce lock-in. Several of these speak the same
systemone-style request format, so swapping providers is cheap. - Does it say what it cannot do? Vendors that publish weak rows get credit.
1. Jev (TypeSafe AI)
Jev started the category. TypeSafe founder Diogo Almeida calls it a "System One Model": it takes text or JSON and returns a choice (one of up to 255 options), a score or a yes/no probability called a noul, each with a calibrated confidence. It cannot generate text.
- Price: $0.042 per million input tokens, output free. New accounts get $5 in credit, which TypeSafe says is roughly 120 million tokens. The waitlist ended on September 21, 2026.
- Speed: TypeSafe reports 70 to 500 ms end to end, against seconds for comparable LLM calls. Its 20 to 200 times faster and 40 to 400 times cheaper claims are TypeSafe's own, and we examined them in our Jev speed and cost fact-check.
- Limits: a 32K context window and, at launch, text and JSON only. InfoQ's summary of TypeSafe's documentation says Jev is unreliable at counting, arithmetic and date comparison.
Why it is first: it has the most field use, from LangSmith trace scoring to the prompt-injection detector use case, and every other entry on this list is measured against it. Start with how Jev works and the launch coverage.
2. pplx-decider v1.1 (Perplexity)
Perplexity open-sourced pplx-decider-v1-27b on October 1, 2026 under Apache 2.0 and launched a hosted Decisions API the same day. It is a multimodal fine-tune of Qwen3.8-27B (about 26B parameters) with a separate decision head and a context window of about 250K tokens, and it takes images as well as text. Perplexity's launch claim was 85.71 percent overall against Jev's 84.51 percent on an 11-benchmark panel.
On October 6 came v1.1. The model card credits lifting the causal mask and more training data, and reports a Decision Index 0.3 overall score of 61.56, up from 56.4 for v1 and ahead of Jev's 57.9. Perplexity's developer account says the hosted price is now $0.02 per million input tokens, half of v1. We could not find that price on the model card, so confirm it on Perplexity's pricing page.
The caveat: gains are uneven. Knowledge scored 48.18 against Jev's 51.4, and the model needs roughly 49 GiB of GPU memory, so self-hosting is not cheap. One outlet noted Perplexity's posts do not include benchmark methodology. Details are in our pplx-decider write-up.
3. OpenAI Decisions API (gpt-6-luna)
OpenAI announced the Decisions API at DevDay on September 30 and moved it to public beta on October 6, 2026. The endpoint is POST /v1/decisions, only gpt-6-luna is supported for now, and it handles three question types: predicate (a probability that a condition is true), choice and score. Images work as inline base64 data URLs.
- Price: $0.10 per million input tokens, with no charge for cache reads, cache writes or output.
- Speed: OpenAI says answers arrive "10x faster than the Responses API." That baseline is OpenAI's own generation endpoint, not Jev.
- Day-one field reports: Hacker News testers measured about 2.4 times Jev's list price, and one measured about 3.1 times on a small eval. Latency was disputed: 160 to 175 ms end to end in one test, a p50 of 346 ms through OpenRouter in another. Simon Willison shipped a plugin the same evening.
Why third: it is the easiest to adopt if you already buy from OpenAI, with zero data retention and HIPAA for eligible customers, and US and EU data residency per the day-one reports. Its case is procurement, not performance. See our DevDay coverage and the day-one developer verdict versus Jev.
4. Cloudflare Clef and Clef-flash
Cloudflare published Clef on October 1, 2026 with Apache 2.0 weights. Clef freezes a Qwen3.8-27B backbone, trains a rank-256 LoRA and adds a routing and classification head, with 64K context and a vision encoder. Clef-flash uses a Qwen3.5-9B backbone for the latency-sensitive hop. Both speak a Jev-compatible API.
Reported prices from early October threads are about $0.24 per million input tokens for Clef and $0.09 for Clef-flash. Cloudflare's own table lists Clef-flash at about 38.8 ms median against Jev at about 524 ms, and a Cloudflare-run index shows Clef leading several rows.
The caveat is the gap between vendor and field. One Hacker News tester reported hosted Clef two to three times slower than Jev and worse on a hate-speech task. Both can be true: a warm median on short prompts is not a distribution. Our decision-models guide walks through the Clef recipe in detail.
5. Liquid AI d1, d1-3B and d1-omni-600M
Liquid AI shipped in two steps. The hosted d1 (October 6) returns probabilities with zero output tokens and now accepts images, at $0.04 per million input tokens. Liquid claims it matches or beats GPT-6.1 Sol on four of six real applications at 19 to 200 times lower cost, and 85 to 97 percent accuracy on four inspection tasks. Model size and weights were not disclosed for that version.
A day later Liquid released open weights: d1-3B (text and images, built on LFM2.5-VL-3B) and d1-omni-600M (text with image or audio, experimental). Liquid reports d1-3B at 16 ms on an NVIDIA Jetson AGX Thor, 48.57 on its Decision Index 0.2.1 and a 82.9 mean over seven public datasets, with the 600M model at 78.4.
Why fifth: it is the best fit for devices and non-text inputs. The caveat is that the cost claim is a vendor comparison against much larger generative models, so measure on your own labeled set. See Liquid d1 and Liquid Open d1.
6. Strands Decider 2B
Published October 1, 2026 by Marc Brooker, Mike Chambers and Fabio Nonato de Paula in the official Strands Agents post, Strands Decider 2B has about 1.9 billion parameters, a Qwen3.5-2B-Base torso, a rank-16 LoRA and a pointer head of roughly 1M parameters. It picks one option from a list or rates something on a scale, with a confidence per answer, and cannot write text. Code and weights are Apache 2.0, and the training data sources and scripts are public.
- Speed: about 115 ms median and 299 ms p95 on an RTX 3090.
- Accuracy: the pinned v19 checkpoint scores 0.723 on the JevBench public set (167 of 231), easy tier 1.000, hard tier 0.505. The repository now treats v21 as the reference at 0.762.
- Training: the authors say you can retrain it in about 11 hours on a single consumer GPU.
Why sixth: it is the most transparent recipe on the list, which matters if you plan to fine-tune. The caveat is the hard tier, where it is near a coin flip. Details in our Strands Decider 2B post.
7. Security-One 27B (Superagent)
Superagent released Security-One on October 5, 2026 as a 27B Apache 2.0 decision model that returns unsafe probabilities for prompts, agent tool calls, code changes and alerts. It is built to run on every event because it generates no tokens.
On BIPIA it caught 599 of 600 attacks at a 0.70 threshold with 1 of 200 benign inputs flagged. On Deepset it caught 47 of 60, about 78 percent, and it showed higher false positives than Jev on NotInject.
The caveat comes from the vendor itself: use it for screening, not enforcement. It belongs in front of a human or a stronger model, behind least privilege and sandboxing. See Security-One 27B.
8. GLiNER2.5-Decide (Fastino)
Fastino announced GLiNER2.5-Decide on September 24, 2026: a 340M-parameter encoder under Apache 2.0 that answers typed questions under logical rules, such as "if safety is safe, harm type must be none," and returns probabilities and confidence. A constrained decoder picks the best joint assignment that satisfies the rules.
Fastino reports a 60.1 percent average across 17 datasets, leading 9, and 167.3 ms p50 at 64 tokens on a 48-vCPU Xeon, so it needs no GPU. A hosted API with fine-tuning is also available.
Why eighth: the rules feature is unique on this list and useful for moderation and compliance logic. The caveat is that the benchmark is Fastino's own suite and the model is small. See GLiNER2.5-Decide.
9. Julia 1 (Supersonic Labs)
Julia 1 (September 27, 2026) is a 144.3M-parameter Apache 2.0 model built on the multilingual mmBERT-small encoder with a decision head, about 550.5 MiB. It runs on a laptop CPU in Python or in a WebGPU browser through an ONNX build, and handles 2 to 20 options per call.
The honest numbers: on one 2,000-item suite it beat a Jev reference by 0.45 points (73.15 percent against 72.70), but on Banking77 it scored 64 percent against 87 on 100 examples. A planned API price of $0.025 per million input tokens was announced, but the API was not live. Julia 1 is one of five models llama.cpp lists for /v1/systemone. See Julia 1.
10. Kev and Jeff (community fine-tunes)
Two independent open families round out the list, and both show how cheap it is to fine-tune this kind of model.
Kev is Jared Palmer's family of 0.8B, 4B and 9B models on Qwen3.5, using a LoRA adapter and a pointer head. Its own README says it trails Jev by about 4.5 points on its development set, yet Kev-9B beat Jev on an independent 900-ticket support test (0.952 against 0.897 on routing). There is no 8B model, despite early headlines. See Kev.
Jeff from firelex has Apache 2.0 0.8B and 2B Qwen3.5 and Gemma fine-tunes with a Jev-shaped local API, about 22 ms on an RTX PRO 6000 and 28 ms on an M4 Max. A held-out voice-navigation task rose from 31.7 percent zero-shot to 95.8 percent after about 30 minutes of fine-tuning on one GPU. The 83.1 percent panel score should not be read as your accuracy. See Jeff.
Why last: these are the fine-tune-it-yourself options, with the least polish and the most control. Others in the same wave include Laya, OpenJev and CLM-8B; our six-clones roundup covers the first wave.
Which one should you pick?
| If you need | Pick |
|---|---|
| The default hosted option for text | Jev |
| Best open weights, multimodal | pplx-decider v1.1 |
| One vendor contract and compliance | OpenAI Decisions API |
| Edge device, image or audio | Liquid d1-3B or d1-omni-600M |
| Retrain on your own labels, small GPU | Strands Decider 2B, Kev or Jeff |
| Prompt-injection screening | Security-One, behind human review |
| Rules between answers, CPU only | GLiNER2.5-Decide |
| Laptop or browser | Julia 1 |
Run decision models locally
llama.cpp added a POST /v1/systemone endpoint in PR #29818 on October 2, 2026. The announcement names five GGUF models: Julia-1, Laya, Kev-4B, lev and OpenJev. Four are Apache 2.0, and OpenJev is CC BY-NC 4.0, which restricts commercial use. Strands Decider's server exposes the same path. If you want a local daemon instead, see Ollaya, and for the llama.cpp specifics read llama.cpp decision model support.
How to test before you switch
Vendor panels disagree, and a 0.45-point lead means nothing on your data. A short protocol:
- Label 200 to 500 real inputs from your own traffic, including ambiguous ones.
- Run three candidates: one hosted, one open model, and your current LLM prompt as the baseline.
- Compare calibration, not only accuracy. Check whether a 0.9 probability is right about 90 percent of the time, since you will threshold on it.
- Measure the tail. Record p95 latency, not the median, on cold and warm calls.
- Route low-confidence cases to your LLM or a human, and count how many that is. A cheap model that escalates 40 percent of traffic saves less than its price suggests.
- Put it behind an adapter so you can swap the base URL later. Our Jev routing integration guide shows the pattern.
What this means for what you build or pay
- Routing and gating hops: move them off a chat model once you have a labeled set. At 300 tokens a call, a million decisions costs roughly $12.60 on Jev, $72 on Clef and about $30 on OpenAI at list price, with output free on all three.
- Do not skip the LLM: decision models cannot write, plan or add options. Keep the LLM for definition and exceptions.
- Do not trust security scores as enforcement. Pair Security-One or Jev with sandboxing, as in our agent sandbox coverage.
- Expect price drops. Perplexity halved its price within a week. Do not lock in annual commitments yet.
- Check licenses. Apache 2.0 for most, CC BY-NC for OpenJev.
What we could not confirm
- Any independent benchmark that ranks all ten on the same data. None exists yet.
- Perplexity's $0.02 price on its model card, and its benchmark methodology.
- Hosted latency for Clef against Jev, where vendor and tester numbers conflict.
- Final pricing and a GA date for OpenAI's Decisions API.
- A live API for Julia 1.
Related reading on explainx.ai
- What are decision models? The practitioner category guide
- TypeSafe AI launches Jev
- Top 10 use cases for Jev
- OpenAI Decisions API developer verdict versus Jev
- Perplexity pplx-decider and the Decisions API
- llama.cpp adds decision model support
- How to wire Jev into your agent pipeline
- Six Jev clones in two days
Prices, benchmarks and availability are vendor-reported as of October 8, 2026 and change quickly. We did not re-run any benchmark.
