deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

OpenAI's Decisions API enters public beta with 10x-faster typed answers

OpenAI's Decisions API is in public beta, returning typed probabilities, choices and rubric scores from gpt-6-luna about 10x faster than the Responses API, with input-only pricing.

OpenAI's Decisions API enters public beta with 10x-faster typed answers

OpenAI has moved its Decisions API into public beta, according to the company's developer documentation, which drew attention on the Hacker News front page. The new endpoint targets a narrow but high-volume job: getting a model to return a judgment rather than generate prose. OpenAI says answers come back roughly ten times faster than through its general-purpose Responses API, and that general availability should follow within weeks.

Three kinds of questions

The service runs on a single model, gpt-6-luna, behind a dedicated POST /v1/decisions endpoint. Each request carries a model, some shared input — either a text string or user messages that mix text and images — and a list of questions. Every question has an explicit type, and there are exactly three.

A predicate question asks whether a condition holds and returns a probability between 0 and 1; the documentation's example inspects a product photo for a crack, tear or dent and gets back a number that can be compared against a review threshold. A choice question picks one value from a caller-supplied set, such as routing a billing complaint to the right department, and returns the chosen value along with a probability distribution over all options and a confidence figure. A score question rates an input against ordered levels — cosmetic, workaround available, fully blocked, for instance — and returns a weighted average of the numbered levels, which can land between two levels when the model is split.

Independent questions can share a single request and a single billable pass over the input, and different types can be mixed within one call. Questions that depend on an earlier answer require separate requests: check for damage first, then classify the repair category only if needed.

Designing for thresholds

Answers arrive in an array, each tied to a name the caller assigns. Because choice and score answers expose the full distribution over options plus a separate confidence field, the intended pattern is threshold engineering rather than binary trust. OpenAI recommends calibrating cut-offs using labeled examples from your own application, weighing the cost of false positives against false negatives, and including a fallback option such as "other" so unmatched inputs flow to a general review queue.

Image evidence is supported, but only as inline base64 data URLs — hosted HTTP image links and file_id references are not accepted by this endpoint.

Pricing and boundaries

gpt-6-luna input is billed at $0.10 per million tokens, and only input is charged: no output tokens, no cache reads, no cache writes. Regional processing premiums and long-context multipliers still apply.

The documentation is also explicit about when not to use the API. Decisions is for cases where the answer is a probability, a category or an ordered rating. When you need a generated object that follows your own JSON schema — extracted fields or a written explanation — OpenAI points to Structured Outputs on the Responses API, and when the model should invoke a tool with arguments, it points to function calling.

Why it matters

A large share of production LLM traffic is classification wearing a generation costume: routing support tickets, triaging bug reports, moderating content, ranking work queues. Doing that through a generative endpoint means paying for output tokens and parsing free text back into labels. A purpose-built endpoint that returns explicit probabilities, choices and rubric scores an order of magnitude faster, with input-only billing at ten cents per million tokens, changes the economics enough to place an LLM judge inside hot request paths — moderation before a post renders, routing before a human ever sees a ticket.

The caveats are real: one model only, no hosted-image input, and a beta label with pricing and behavior that can shift before general availability. But the shape of the release — typed outputs, per-option probabilities, confidence values, threshold guidance — signals that OpenAI now treats "model as classifier" as a first-class product rather than a prompt-engineering workaround, and that is a useful marker for where inference infrastructure is heading.

  • #openai
  • #api
  • #llm
  • #classification
  • #developer-tools

Related posts