· via Hacker News – Front Page (native)
OpenAI's Decisions API enters public beta with 10x-faster typed answers
OpenAI's Decisions API is in public beta, returning typed probabilities, choices and rubric scores from gpt-6-luna about 10x faster than the Responses API, with input-only pricing.

OpenAI has moved its Decisions API into public beta, according to the company's developer documentation, which drew attention on the Hacker News front page. The new endpoint targets a narrow but high-volume job: getting a model to return a judgment rather than generate prose. OpenAI says answers come back roughly ten times faster than through its general-purpose Responses API, and that general availability should follow within weeks.
Three kinds of questions
The service runs on a single model, gpt-6-luna, behind a dedicated POST /v1/decisions endpoint. Each request carries a model, some shared input — either a text string or user messages that mix text and images — and a list of questions. Every question has an explicit type, and there are exactly three.
A predicate question asks whether a condition holds and returns a probability between 0 and 1; the documentation's example inspects a product photo for a crack, tear or dent and gets back a number that can be compared against a review threshold. A choice question picks one value from a caller-supplied set, such as routing a billing complaint to the right department, and returns the chosen value along with a probability distribution over all options and a confidence figure. A score question rates an input against ordered levels — cosmetic, workaround available, fully blocked, for instance — and returns a weighted average of the numbered levels, which can land between two levels when the model is split.
Independent questions can share a single request and a single billable pass over the input, and different types can be mixed within one call. Questions that depend on an earlier answer require separate requests: check for damage first, then classify the repair category only if needed.
Designing for thresholds
Answers arrive in an array, each tied to a name the caller assigns. Because choice and score answers expose the full distribution over options plus a separate confidence field, the intended pattern is threshold engineering rather than binary trust. OpenAI recommends calibrating cut-offs using labeled examples from your own application, weighing the cost of false positives against false negatives, and including a fallback option such as "other" so unmatched inputs flow to a general review queue.
Image evidence is supported, but only as inline base64 data URLs — hosted HTTP image links and file_id references are not accepted by this endpoint.
Pricing and boundaries
gpt-6-luna input is billed at $0.10 per million tokens, and only input is charged: no output tokens, no cache reads, no cache writes. Regional processing premiums and long-context multipliers still apply.
The documentation is also explicit about when not to use the API. Decisions is for cases where the answer is a probability, a category or an ordered rating. When you need a generated object that follows your own JSON schema — extracted fields or a written explanation — OpenAI points to Structured Outputs on the Responses API, and when the model should invoke a tool with arguments, it points to function calling.
Why it matters
A large share of production LLM traffic is classification wearing a generation costume: routing support tickets, triaging bug reports, moderating content, ranking work queues. Doing that through a generative endpoint means paying for output tokens and parsing free text back into labels. A purpose-built endpoint that returns explicit probabilities, choices and rubric scores an order of magnitude faster, with input-only billing at ten cents per million tokens, changes the economics enough to place an LLM judge inside hot request paths — moderation before a post renders, routing before a human ever sees a ticket.
The caveats are real: one model only, no hosted-image input, and a beta label with pricing and behavior that can shift before general availability. But the shape of the release — typed outputs, per-option probabilities, confidence values, threshold guidance — signals that OpenAI now treats "model as classifier" as a first-class product rather than a prompt-engineering workaround, and that is a useful marker for where inference infrastructure is heading.
- #openai
- #api
- #llm
- #classification
- #developer-tools