OpenAI released the Decisions API in public beta on 6 October 2026. It accepts text, images or both and returns one of three typed answers: a probability for a condition, a choice from an allowed set, or a score against ordered levels. OpenAI says it is about ten times faster than the Responses API and currently supports only `gpt-6-luna`. That makes it useful for classification, routing and prioritization—but it does not make a probability an authorization decision. Production value comes from calibrated thresholds, an explicit review lane, deterministic policy checks and an audit trail around every answer.
The official changelog and Decisions guide were verified on 7 October 2026. The launch date, beta status, supported model, question types, pricing and data-control statements below come from OpenAI. Thresholds, service boundaries and rollout steps are engineering recommendations, not OpenAI guarantees.
What OpenAI actually released
The dedicated endpoint is `POST /v1/decisions`. A request supplies `model`, shared `input`, and a `questions` array. The input may be a text string or user messages containing text and images. Each question has a unique name, a type and instructions; the response echoes the name in an `answers` array so application code can map results without parsing prose.
There are three contracts. A `predicate` estimates the probability that a condition is true. A `choice` selects one value from a fixed, unordered set and includes probabilities over the available choices. A `score` evaluates ordered levels and returns a probability-weighted average of their indices. OpenAI positions these for tasks such as visible-damage checks, department routing, answer relevance and severity scoring.
The API is a public beta, not a stable replacement for every Responses workflow. OpenAI says `gpt-6-luna` is the only available model and expects general availability in the coming weeks. Treat both the endpoint and SDK surface as versioned dependencies: pin supported SDK versions, add contract tests, and keep a reversible path until the beta proves stable for your traffic.
Choose the right primitive before writing a prompt
The Decisions API is narrow by design. Use it when the output is exactly a probability, a closed-set category or an ordered score. Use Structured Outputs with the Responses API when you need a custom JSON object, extracted fields or a written explanation. Use function calling when the model must propose a tool and arguments. Use ordinary code, SQL or rules when the answer already exists deterministically in your system.
Need Best starting point
Boolean-like condition + probability Decisions predicate
Closed, unordered category Decisions choice
Ordered rubric or severity Decisions score
Custom JSON extraction or explanation Responses + Structured Outputs
Action request with tool arguments Function calling + authorization
Exact entitlement, balance or policy Code / SQL / rules engine
Search for matching records Search or vector retrievalA common anti-pattern is asking a model to decide whether a refund is allowed when the policy is already encoded in transaction data. Query the order, payment and policy tables instead. Another is using a score to extract names, dates and line items; that is a structured extraction problem. Choosing the narrowest correct primitive reduces ambiguity, latency and operational risk.
Separate inference from policy
The model answer is evidence for a decision, not the whole decision. Put a policy service after the API. It validates the answer shape, joins trusted application facts, applies thresholds for the current use case, and chooses an action such as auto-route, request more evidence, send to review or decline automation. Permissions and irreversible effects stay outside the model.
A practical flow has six stages: validate the incoming evidence; normalize it without erasing useful detail; call Decisions with a versioned question set; evaluate the typed answer against a versioned policy; execute only an authorized and idempotent action; and record the evidence reference, question version, model, distribution, threshold, policy result and final human outcome. The audit log should let you reproduce why the system routed a case without storing unnecessary sensitive content.
const result = await client.decisions.create({
model: "gpt-6-luna",
input: ticketText,
questions: [{
type: "choice",
name: "queue",
instructions: "Route by the customer's requested outcome.",
choices: ["billing", "delivery", "technical", "other"]
}]
});
const answer = result.answers.find(x => x.name === "queue");
const policy = routeWithReview({
answer,
questionVersion: "support-queue-v3",
autoRouteMinConfidence: 0.92,
highImpact: ticketRequestsAccountClosure(ticketText)
});The threshold and helper names are illustrative application policy, not fields promised by OpenAI. The important boundary is that the service may recommend a route, while application code owns the threshold, account permissions, side effects and recovery.
Calibrate thresholds against the cost of mistakes
A raw probability is not automatically calibrated for your population. A value of 0.90 does not prove that nine of ten production cases are correct. Build a labeled set that resembles real traffic, then compare predictions with known outcomes. Segment by language, channel, image quality, customer group and input length because one global average can hide a weak cohort.
Set thresholds from business loss, not convenience. For a low-impact inbox route, a false positive may cost one internal reassignment. For visible product damage, a false negative may deny a legitimate claim. For fraud, safety or eligibility, both errors may be expensive and legally sensitive. Define separate auto-accept, review and auto-reject regions rather than forcing every score through one boundary.
Track precision, recall, false-positive rate, false-negative rate, review rate, coverage, calibration error, p50 and p95 latency, and cost per correctly automated case. OpenAI's guide explicitly recommends labeled application examples and thresholds based on the cost of false positives and false negatives. Recalculate after changing instructions, choices, score levels, input preprocessing or model behavior.
Design questions that remain observable
Ask about evidence visible in the input. “Is there a dent in the product, excluding packaging and shadows?” is testable. “Is this customer trustworthy?” compresses identity, history and social judgment into an opaque label. When a question depends on trusted account state, fetch that state separately and let code combine it with the model output.
Give categories distinct meanings and include an `other` or review path when the world is broader than the taxonomy. For scores, define adjacent levels with concrete boundaries so reviewers can label them consistently. Do not hide three independent concerns inside one question; damage, category and urgency should be separate named answers.
The API can evaluate multiple independent questions over shared input in one request. OpenAI recommends separate requests when one decision depends on an earlier answer. Preserve that distinction. A conditional second step can use the validated first result and a narrower question, while a single overloaded request makes dependencies harder to test and audit.
Build an uncertainty and review lane
Production systems need a valid outcome for uncertainty. Send low-confidence answers, close probability ties, unsupported categories, malformed inputs and high-impact cases to review. Show reviewers the original evidence, the typed distribution, the question and policy versions, and the proposed action. Let them correct both the outcome and a structured reason. Those corrections become evaluation data—not silent overrides that disappear.
OpenAI's safety guidance recommends adversarial testing and human review, especially in high-stakes domains. Review is not a failure of automation. It is how the system limits harm while collecting the exact edge cases needed to improve coverage safely.
Do not let a probability bypass authentication, ownership checks, approval requirements or spending limits. A “safe” classification should never expand authority. If the result triggers a mutation, use an idempotency key and a durable state transition so retries cannot repeat refunds, messages or account changes.
Evaluate the complete decision system
Offline evaluation should freeze the question version, model, policy and dataset snapshot. Include typical traffic, rare but costly cases, ambiguous inputs, blurry images, multilingual examples, missing context and adversarial instructions embedded in evidence. Compare the full route—not only whether the top category matched. A useful scorecard asks whether the correct cases automated, the uncertain cases reached review, and forbidden actions remained impossible.
OpenAI's evaluation guidance recommends task-specific datasets that reflect production distributions, continuous evaluation, and calibration of automated metrics with human judgment. For Decisions, maintain both a held-out release set and a rolling sample of production outcomes. Do not tune thresholds on the same records used to report success.
During a shadow phase, call the beta endpoint but keep the existing system authoritative. Compare its route, latency and review demand with the current path. Then canary by a stable cohort, monitor error costs and review queues, and define rollback thresholds before increasing exposure. A beta endpoint should not become a single point of failure without timeout, circuit breaker and deterministic fallback behavior.
Account for latency, price and data boundaries
OpenAI states that Decisions is about ten times faster than the Responses API. Treat that as a provider claim, then measure your end-to-end path: upload time for images, request queueing, inference, policy lookup, human-review wait and downstream execution. A faster classifier cannot fix slow evidence storage or an overloaded review queue.
For `gpt-6-luna`, OpenAI lists input pricing at $0.10 per million tokens for `/v1/decisions`, with no cache-read, cache-write or output-token charges. Regional processing premiums and long-context multipliers may still apply. Measure cost per correct automated decision, not only cost per request; uncertain or wrong results create review, rework and customer-support costs.
The official guide says the API supports Zero Data Retention and HIPAA use for eligible customers, plus data residency and regional processing in the United States and Europe (EEA and Switzerland). Eligibility, agreements and limitations still matter. Minimize inputs, avoid sending fields that do not affect the question, apply tenant isolation before the call, and set retention and access controls for your own logs.
When not to use Decisions
Do not use it where a database lookup, rule or calculation produces the authoritative answer. Do not use it to generate explanations, summaries or arbitrary JSON. Do not let it make unreviewed medical, legal, employment, credit, safety or other high-impact determinations. Do not use a visual predicate when the evidence requires measurement the image cannot support. And do not automate a route when your labels are inconsistent or your review team cannot explain the rubric.
Avoid threshold theater: a number with two decimals can look precise while the underlying task, dataset and model remain uncertain. Avoid silently changing instructions because that changes the classifier contract. Avoid using the top probability alone when the leading choices are nearly tied. Avoid logging raw customer evidence by default just because an audit trail is needed; references, hashes and minimal derived facts may be enough.
If the product needs a nuanced explanation or tool-using workflow, route the case to the Responses API or a human after the initial decision. The new endpoint is valuable because it is narrow, not because it replaces every AI architecture.
Executive rollout checklist
Start with one low-impact, high-volume route whose labels are already understood. Write observable questions and a versioned taxonomy. Build a representative labeled set. Choose thresholds from error costs. Add a review region and deterministic permission checks. Shadow the current workflow. Canary a stable cohort. Monitor calibration, review capacity, latency and cost per correct automation. Keep a circuit breaker and fallback. Promote only when the whole system—not a demo request—meets the product's acceptance criteria.
The Decisions API can reduce the overhead of turning multimodal evidence into a small typed result. Its durable advantage will come from the surrounding engineering: clear contracts, calibrated policy, constrained authority, human correction and evidence that the automated route is better than the process it replaces.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




