OpenAI's Decisions API went into public beta on October 6, 2026, and within about two hours developers had pointed it at their own workloads. The Hacker News thread (roughly 190 points and 80+ comments) is the first real field report. One tester ran around 600 calls against Jev and Mercury Decide, another timed 160 to 175 ms end to end, and Simon Willison shipped a plugin before the evening was over.
The short verdict from those reports: POST /v1/decisions with gpt-6-luna is cheap and quick in absolute terms, but on day one it did not beat Jev on list price, on one tester's measured cost, or on confidence for ambiguous inputs. Its strongest case is not performance. It is procurement, compliance, and not adding a vendor. Every number below is a commenter's self-report from a small, launch-day sample, not a verified benchmark.
This post does not re-explain the announcement or the endpoint table. Those are in the DevDay write-up and its public beta update. Here the question is what happened when people actually used it.

TL;DR: the day-one verdict
| Question | Day-one answer |
|---|---|
| Is it cheaper than Jev? | No on list price: $0.10 vs $0.042 per million input tokens, about 2.4x. One tester measured about 3.1x on a small eval. Output is free on both. |
| Is it faster? | Disputed. One tester saw p50 346 ms and p95 860 ms via OpenRouter and called it slower than Jev. Another saw 160 to 175 ms end to end and p95 285 ms. |
| Is it better calibrated? | Nobody knows. OpenAI makes no calibration claim in the docs we read. One tester saw lower probabilities than Jev on ambiguous tags. |
| Does it take images? | Yes, as inline base64 only. Jev was text and JSON only at launch, which commenters called a gap. |
| Why pick it anyway? | One vendor contract, ZDR and HIPAA for eligible customers, US and EU residency, and an SDK you already have. |
Will /v1/decisions become a standard? | Simon Willison thinks it might. Other vendors' schemas still differ today. |
| Should I switch production traffic? | Not on day one. Add it as a second provider behind an adapter and run your own eval. |
What developers on Hacker News are saying
Simon Willison: a working call, a plugin, and a standards bet
simonw posted a curl call with two independent predicate questions about the sentence "I am angry about the new product feature." The API returned a probability of 0.91 for "is this a complaint" and 0.06 for "is this a compliment." His pasted usage block showed 310 input tokens and zero output tokens, which matches the docs' no-output-charge pricing.
His bigger point was about the endpoint itself. When OpenAI defines an endpoint shape like /v1/decisions, he wrote, it often becomes a de facto standard that other providers copy. That is a prediction, not a fact yet. Perplexity's decider uses choice and noul, Jev uses state plus typed questions, and llama.cpp serves /v1/systemone, as covered in our llama.cpp post. Whether those converge on OpenAI's shape is open.
He also released a plugin the same evening. We checked it on GitHub: simonw/llm-openai-decisions was created on October 6, is Apache-2.0, and has a 0.1 release on PyPI. The README documents predicates, choices, scores, multiple named questions per request, and refusals preserved as refusal answers. It accepts PNG, JPEG, WebP, and GIF attachments from a URL or local path. The source builds data: base64 URLs from those attachments, so the plugin hides the inline-base64 requirement. Install is llm install llm-openai-decisions, then llm keys set openai.
In a second comment, simonw made the point the rest of this post leans on: quality will decide the market, and anyone using a decision model has to build their own evals, because these outputs are much harder to vibe-check than text.
Topfi's roughly 600-call eval against Jev and Mercury Decide
The most detailed report came from Topfi, who ran fewer than 600 calls across UI component selection, chat charting, tag selection, and personal-knowledge-management tasks, through OpenRouter, against Jev and Mercury Decide. His own caveats are worth repeating: the evals are rudimentary, the use cases are unusual (a PKM-focused Firefox fork with infinite canvases), and he had not yet tested image input.
What he reported for Luna:
- Latency: p50 346 ms and p95 860 ms, slower than Jev and roughly on par with Mercury Decide, though Luna's latency did not grow linearly with input size.
- Confidence: lower confidence on his ambiguous UI-component and response-shape tasks, often at 0.6 and below.
- Failures: four failed calls, where neither competitor had any.
- Cost: about 3.1x Jev on average.
His summary was blunt: slower, pricier, and less capable than Jev for his workload, roughly on par with Mercury Decide, and "a bit undercooked." He added that he would rather frontier labs not jump on bandwagons without a price or performance edge. A single tester on one workload is a data point, not a ruling. With a sample this small and OpenRouter in the path, four failures cannot be separated into preview-day errors, routing-layer errors, or a systematic issue.
Latency: two testers, two pictures
agentdev001 reported a very different picture: 160 to 175 ms end to end, the worst 5 percent at 285 ms, and the worst single call at 743 ms. They also noted that because this is gpt-6-luna underneath, it should offer a 1M-token input window and image input. That window is the commenter's inference. The docs pages we read do not state a context limit.
| Source | Reported latency | Setup disclosed |
|---|---|---|
| Topfi | p50 346 ms, p95 860 ms | Via OpenRouter, mixed tasks, fewer than 600 calls across three models |
| agentdev001 | 160 to 175 ms typical, p95 285 ms, worst 743 ms | End to end, route and sample size not stated |
| OpenAI docs | "10x faster than the Responses API" | Relative claim, no absolute figure |
| OpenAI forum post | "up to 10x faster than GPT-6 Luna through the Responses API" | Relative claim with an "up to" |
These are not necessarily contradictory. Our hypotheses, none verified: an OpenRouter hop adds time that a direct call does not; Topfi's task-specific inputs may be longer than a one-line complaint; launch-day preview load varies (Topfi himself blamed preview conditions for Mercury Decide's spikes); and neither tester published a methodology we can replicate. The Jev side is also fuzzy. TypeSafe's own materials describe a roughly 100 ms class, while the Cloudflare table cited in our decision-model category guide listed Jev near 524 ms. "Slower than Jev" depends on which Jev number you hold in your head. Another commenter, lelandfe, noted that self-hosting removes network round trips entirely, which no hosted API can match.
Price: list rate versus measured cost
On list price the answer is simple. jerrygenser put it at $0.10 per million input tokens versus $0.042 for Jev, with free output on both. Topfi followed up that in the same bench a full Jev run cost $0.0192 against $0.06 for Luna, both via OpenRouter, about 3x in Jev's favor.
The ratio deserves a second look. The list-price gap is 2.4x, but Topfi measured about 3.1x. If OpenRouter charged list rates, that would imply Luna billed roughly 30 percent more input tokens for the same eval. That is our arithmetic, not Topfi's, and it may reflect OpenAI's per-request scaffolding. Simon's 310-token usage for a roughly ten-word input hints at a fixed overhead of several hundred tokens per request, again our inference from one data point.
Scale decides whether any of this matters. At 300 tokens per call, one million decisions cost about this much:
| Hosted option | Input price per 1M tokens | Cost of 1M decisions at 300 tokens |
|---|---|---|
OpenAI Decisions API (gpt-6-luna) | $0.10 | $30.00 |
| Jev | $0.042 | $12.60 |
| Perplexity Decisions API | $0.04 (a cut to $0.02 was reported by aggregators, unconfirmed) | $12.00 |
| Liquid d1 | $0.04 | $12.00 |
| Clef-flash | about $0.09 | about $27 |
| Clef | about $0.24 | about $72 |
Rates for the non-OpenAI rows come from our posts on Jev's launch, Perplexity's decider, Liquid d1, and the category guide. At 10 million decisions a month the Luna-versus-Jev gap is about $174. At 100 million it is about $1,740. For one commenter, tmhall, the whole workload comes to roughly $11 a month, and they already have OpenAI billing in place, so the cheaper vendor is not worth a new contract.
Calibration: the question nobody could answer yet
AM1010101 asked the question that matters most for anyone thresholding on these numbers: is a 90 percent output right nine times in ten? They noted that Jev had put visible effort into calibration and that it is unclear whether Luna is calibrated the same way. TypeSafe does make an explicit calibration claim for Jev's training, but our fact-check of Jev's speed and cost claims found independent verification still thin. OpenAI makes no calibration claim at all in the pages we read, only that choices and scores return a confidence and per-option probabilities.
Topfi's tag tests are the only calibration-flavored evidence so far. Each task asked whether to apply a tag to a browser tile title, with a 60 percent threshold, and every task was run twice per model:
| Tile title | Tag | Jev (two runs) | Luna (two runs) |
|---|---|---|---|
| Mortgage calculator, Bankrate | house hunting | 0.69, 0.72 (yes) | 0.21, 0.21 (no) |
| S&P 500 index live chart, Bloomberg | investing | 0.91, 0.90 (yes) | 0.56, 0.56 (no) |
Topfi's read is that Jev's values matched his own judgment better, and that Luna's output was not useful for his goal of an instant, overridable default tag. He also granted that tags are subjective and that results vary by task. Two observations from the table are ours: Luna's repeated runs matched to two decimals while Jev's drifted by a point or three, and the S&P 0.56 sits just under the 60 percent cutoff. The practical lesson is that a threshold tuned on one model does not transfer to another. A decision model's 0.56 is only meaningful against labeled examples for that model. For the broader method, see our guide to calibrating LLM classifiers.
Images: Jev's gap, not Luna's edge
mritchie712 said image input was the first big gap they found in Jev, and stillatit agreed that Decisions can take images where Jev, for now, cannot. Our own Jev launch coverage confirms text and JSON only at launch, with a 32K window.
The gap is Jev-specific rather than a Luna exclusive. Liquid d1, pplx-decider, and Clef all take images, as our posts on d1, pplx-decider, and the category guide describe. What is distinctive about OpenAI's version is the constraint: images must be inline base64 data URLs. Hosted URLs and file_id inputs are unsupported. You download, encode, and ship the bytes with every call. The docs pages we read do not say how images are tokenized, so measure usage.input_tokens on your real images before estimating cost. Liquid, by comparison, publishes a patch-based rate.
Compliance and the existing-vendor argument
The most consistent defense of Luna had nothing to do with benchmarks. Shank argued that customers with compliance needs will pick OpenAI because they cannot pick Jev. csharpminor said an enterprise with an existing OpenAI procurement agreement avoids onboarding another vendor. jcims described three- to six-month vendor onboarding cycles plus ongoing governance burden. oh_no agreed it makes sense even if underbaked, and added that it is a stopgap: stand it up, see the value, and switch in a few months if something better appears.
The official docs back the factual core. The Decisions API supports Zero Data Retention and HIPAA use "for eligible customers," with data residency and regional processing in the United States and Europe (EEA plus Switzerland). Two cautions: "eligible" depends on your agreement with OpenAI, and regional processing premiums apply on top of the $0.10 rate, so the EU price is not the headline price. We did not verify TypeSafe's compliance posture, so Shank's claim about Jev stays an assertion.
The skeptics: coherence, drift, caching, and novel domains
Coherence across questions. OutOfHere argued that decisions only suit very simple problems, and that making twenty independent calls loses coherence compared with one structured output covering many attributes. The docs half-agree: multiple independent questions can share one request, but dependent decisions need separate requests. Simon's own example shows the issue in miniature, since two independent predicates could both score high. For mutually exclusive categories, use one choice question, whose probabilities form a single distribution (the docs' sample sums to 1.0). Enforce cross-question rules in your own code. devin countered that people want more determinism, and this is a step toward it.
Silent model changes. drdexebtjl did not want a decision model to change underneath them when the lab decides to improve or "make it safer." The docs we read show one model alias, gpt-6-luna, and Simon's response echoed "model": "gpt-6-luna" with no dated snapshot visible. We could not confirm whether version pinning exists. Treat it as a question for OpenAI and defend yourself with a canary set (see below).
Caching. AM1010101 wondered why caching was skipped, since a long shared system prompt would save money on bulk work. BoorishBears guessed it chases the lowest possible latency. That guess is speculation. The docs say only that there are no cache charges.
Novel domains. nostrebored said Jev had failed to understand any novel domain they tried, and would be shocked if Luna were less generally capable. Topfi replied that results are task-dependent and his measurements had not shown Luna ahead yet. That is a genuine split, and we found no published head-to-head of Luna Decisions against Jev on a shared panel. Perplexity's 11-benchmark table shows Jev winning hard reasoning rows like BBH and WinoGrande, but it did not include Luna.
The moat debate
TSiege called OpenAI's response the nail in the coffin for the idea that AI is not a commodity market, arguing cheap System One models start a price war and the big labs now race to make products sticky. dvt saw "absolutely zero moat" and found it hard to see why anyone would pay a premium over Jev or open models. mediaman answered with the rent-versus-buy analogy: at $0.10 per million, convenience beats self-hosting even when Jev is cheaper. TSiege added that the real moat is payments, platform, and branding. scosman noted that planting a flag now and shipping a better model in a few weeks is a rational move even if today's evals are worse.
Our read: both camps are right. The vendors are separated by single-digit multiples on tiny per-call prices. The differentiators are the contract, the SDK, and the compliance paperwork.
What people are asking
Can I use my ChatGPT subscription? No. OutOfHere pointed out that ChatGPT subscriptions have never covered API calls; Decisions bills as API usage.
Is this just Luna in a wrapper? ashu1461 reasoned that the input price matches Luna's, so the difference is speed. rockinghigh added that output is free here. Per our Luna pricing post, Luna on the standard API is $0.10 per million input and $0.50 per million output, so the savings are the output tokens and the latency. Since classification outputs are short, ashu1461 countered that the output saving is small.
Will this fold into the main models? swader999 asked why routing is not simply built into every model, and jasonjmcghee wondered whether it will fold into post-training and improve calibrated outputs. Nobody has an answer.
Do I need new HTTP code? No. OutOfHere noted that the Python SDK covers it. The docs require Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, or Java 4.78.0 and later.
How to call it (from the official docs)
This is OpenAI's own routing example, lightly trimmed. You supply model, input, and a list of uniquely named questions. A choice question takes choices, each with a value and a description.
from openai import OpenAI
client = OpenAI()
decision = client.decisions.create(
model="gpt-6-luna",
input="I was charged twice for my order.",
questions=[
{
"type": "choice",
"name": "department",
"instructions": "Which department should handle this complaint?",
"choices": [
{"value": "billing", "description": "Payments, invoices, and refunds."},
{"value": "technical", "description": "Problems using the product."},
{"value": "shipping", "description": "Delivery and tracking."},
{"value": "other", "description": "Requests outside these categories."},
],
}
],
)
answer = decision.answers[0]
if answer.type == "refusal":
print(f"Refused: {answer.name}")
elif answer.type == "choice":
print(f"Department: {answer.choice} (confidence: {answer.confidence})")
The docs' sample response gives billing at 0.95 with an overall confidence of 0.93, plus a probability for every option. The other two types follow the same pattern. A predicate returns a probability from 0 to 1. A score question takes ordered levels and returns a probability-weighted average of the level indices, so three levels give a range of 0 to 2.
A few rules from the docs that matter more than they look:
- Use distinct choice values with descriptions that say when each applies, and include a fallback such as
other. - Write questions around observable criteria and separate different concerns into different questions.
- Handle
refusalanswers explicitly. The docs show them as a valid answer type. - Send dependent decisions as separate requests.
- For images, use
input_imagecontent with adata:image/png;base64,...URL. Nothing else is accepted.
For quick experiments from a shell, the llm plugin from the HN thread is the fastest path:
llm install llm-openai-decisions
llm -m openai-decisions/gpt-6-luna 'I was charged twice.' \
-s 'Which department should handle this message?' \
-o answer_type choice \
-o choices '{"billing": "Charges, invoices, and refunds", "technical": "Problems using the product", "other": null}'
For the same hop wired into an agent loop, see how to integrate Jev for agent routing and the browser version in Jev Ultrafast. The adapter pattern is the same.
When to choose Decisions API, Jev, or open weights
| Situation | Lean toward | Why |
|---|---|---|
| Already on an OpenAI contract, with ZDR, HIPAA, or EU residency needs | Decisions API | The paperwork already exists, and docs list these controls for eligible customers |
| Image or screenshot inputs | Decisions API, or another vision-capable decider | Jev was text and JSON only at launch |
| Low volume (under about 10M calls a month) | Whichever you can integrate fastest | The list-price gap is a rounding error |
| High volume where cost per decision dominates | Jev or a cheaper hosted decider | Roughly 2.4x gap on list price, 3.1x in one measured run |
| You must pin a model version or keep data in your own network | Open weights such as pplx-decider or the llama.cpp models | You control the checkpoint, though large ones need real GPUs |
| Ambiguous labels where confidence decides routing | Whichever wins on your reliability plot | Day-one evidence is one tester's two tag examples |
| Unsure | Both, behind an adapter | Fail over on confidence, not on brand |
Open weights are not free. Perplexity's 27B decider needs roughly 49 GiB of GPU memory per our pplx-decider coverage, while the smallest llama.cpp models are in the hundreds of millions of parameters.
Run your own eval before switching
Simon Willison said anyone using a decision model has to build their own evals, and Topfi's thread is that advice in practice. Here is a minimal shape for it.
- Label real inputs. Collect 200 or more from your traffic, deliberately including ambiguous ones. Topfi's most informative cases were the arguable ones.
- Run each input at least twice per model. Topfi did, and it exposed that one model repeated to two decimals while another drifted.
- Record failures and refusals separately from wrong answers. A failed call is a different problem.
- Measure latency from your own region with enough calls for a stable p95, on more than one day. Launch-day preview load is not steady state.
- Log billed input tokens, especially for images, and compute dollars per thousand decisions.
- Fit the threshold per model on a validation split, then test on held-out data. Never reuse Jev's cutoff for Luna or the reverse.
- Keep a canary set of 50 to 100 frozen examples. Re-run it daily and alert when probabilities shift, which is your defense against the silent-update worry.
- Compare reliability, not just accuracy. Bucket predictions by probability and check whether 0.9 means about 90 percent. Our guide to reading AI benchmarks and post on evaluating prompts cover the habits.
Here is a harness that does the measuring for the predicate case. We exercised its reporting code against a stub client, but we have not pointed it at the live endpoint, so treat it as a starting point and expect to adapt field names to the SDK you install.
import json, statistics, sys, time
from openai import OpenAI
client = OpenAI()
QUESTION = "Is this page a good fit for the tag 'investing'?"
RUNS = 2 # repeat every row to see run-to-run drift
def call_luna(text):
start = time.perf_counter()
try:
d = client.decisions.create(
model="gpt-6-luna",
input=text,
questions=[{"type": "predicate", "name": "fit", "instructions": QUESTION}],
)
except Exception as exc:
return {"status": "error:" + type(exc).__name__, "ms": (time.perf_counter() - start) * 1000}
ms = (time.perf_counter() - start) * 1000
a = d.answers[0]
if a.type == "refusal":
return {"status": "refusal", "ms": ms}
usage = getattr(d, "usage", None)
return {"status": "ok", "ms": ms, "p": a.probability, "tokens": getattr(usage, "input_tokens", 0)}
def pct(values, q):
s = sorted(values)
return s[min(len(s) - 1, int(q * len(s)))]
def report(results):
ok = [r for r in results if r["status"] == "ok"]
print(f"calls={len(results)} failed_or_refused={len(results) - len(ok)}")
ms = [r["ms"] for r in ok]
print(f"p50={pct(ms, 0.5):.0f}ms p95={pct(ms, 0.95):.0f}ms")
print(f"brier={statistics.fmean((r['p'] - r['label']) ** 2 for r in ok):.4f}")
print(f"input cost at $0.10/M: ${sum(r['tokens'] for r in ok) / 1e6 * 0.10:.4f}")
for i in range(10): # reliability: does 0.9 mean about 90 percent?
lo, hi = i / 10, (i + 1) / 10
b = [r for r in ok if lo <= r["p"] < hi or (i == 9 and r["p"] == 1.0)]
if b:
print(f"{lo:.1f}-{hi:.1f} n={len(b):3d} predicted={statistics.fmean(r['p'] for r in b):.2f} "
f"actual={statistics.fmean(r['label'] for r in b):.2f}")
if __name__ == "__main__":
rows = [json.loads(line) for line in open(sys.argv[1])] # {"text": ..., "label": 0 or 1}
results = [{**call_luna(r["text"]), "label": r["label"]} for r in rows for _ in range(RUNS)]
report(results)
Point the same loop at Jev or an open model by swapping call_luna for another adapter, keep the labels fixed, and compare the reports side by side.
Honest limits of this verdict
- Every figure is self-reported. Topfi, agentdev001, and the others are single users with small samples and unpublished methods. We ran no benchmark of our own.
- Topfi's workload is unusual, and by his own account he had not tested images. A tagging or UI-selection result may not predict your ticket router.
- The latency conflict is unresolved. We offered hypotheses, not findings.
- No head-to-head accuracy table exists yet for Luna Decisions against Jev on a shared benchmark panel, as far as we found.
- Docs gaps: the pages we read give no context limit, image token rates, question or choice limits, or version-pinning policy. The 1M window and the 128-image cap come from a commenter and a plugin README respectively, not from OpenAI.
- Beta pricing and behavior can change. OpenAI says general availability is expected in the coming weeks.
- HN timing: the thread was created late on October 6 UTC and was still active on October 7, so some reports may be updated after we read them.
What this means for what you build or pay
If you already ship a decision hop, add Luna as a second provider behind your adapter, mirror a slice of traffic, and compare reliability and dollars per correct decision on your own labels. If your organization needs OpenAI's compliance controls, the day-one numbers do not argue against it, because the price gap is small money at normal volumes. If cost at scale or a pinned model is the constraint, the early evidence favors Jev or open weights, but not conclusively.
Do not threshold on probabilities you have not checked. The surest lesson from day one is Simon Willison's: these models are hard to vibe-check, so the eval is the product decision.
Related on explainx.ai
- OpenAI Decisions API: GPT-6 Luna at DevDay, plus the public beta update
- What are decision models? The practitioner category guide
- Perplexity pplx-decider and its Decisions API
- How to integrate Jev for agent routing
- Jev Ultrafast: browser-use with typed decisions
- llama.cpp adds decision model support
- Is Jev's speed and cost claim true?
- Stop using LLMs as classifiers: calibration guide
- Official: Decisions API docs · OpenAI community announcement · Hacker News thread · llm-openai-decisions on GitHub
Figures come from OpenAI's Decisions API documentation and community announcement, the Hacker News thread "Decisions API is in public beta" (created October 6, 2026 UTC), and the llm-openai-decisions repository, as of October 7, 2026. Commenter numbers are their own reports and were not independently verified. The public beta may change pricing, limits, and behavior. Run your own evals before you ship.
