The OpenAI Decisions API entered public beta on October 6, 2026, powered exclusively by the gpt-6-luna model and accessible via the dedicated POST /v1/decisions endpoint, running up to 10 times faster than the Responses API by returning typed answers directly without token-by-token text generation.
Why GPT-6 Luna Skips Token Generation
Large language models traditionally burn computational cycles generating text word by word, even when an application merely needs a binary judgment or a categorical routing choice. developers.openai.com reported that the Decisions API cuts out this unnecessary generation step entirely. Instead of forcing a model to write out a reasoning chain and a prose response, the endpoint evaluates text, images, or both, and immediately returns structured, typed data.
The performance gain comes from eliminating the generation bottleneck, not merely scaling underlying hardware compute. For high-throughput infrastructure processing tens of thousands of requests daily, removing output token generation changes the fundamental economics of automated classification.
The API handles requests through three distinct question types. Predicate questions evaluate conditions—such as checking a product photo for visible damage—and return a probability estimation between 0 and 1. Choice questions select a single option from an unordered list, returning probabilities and a confidence score for categories like billing, technical support, or account access. Score questions rate inputs against ordered levels, calculating a probability-weighted average of numeric level indices that can fall between discrete steps, such as severity ratings.
Pricing Structures and Enterprise Token Costs
The financial model of the Decisions API mirrors infrastructure utilities rather than conversational chat interfaces. Input tokens cost $0.10 per million tokens. Developers pay exclusively for input tokens, with zero cache-read, cache-write, or output-token charges. Regional processing premiums and long-context input pricing multipliers apply for eligible deployments.
By abandoning output-token billing for this endpoint, OpenAI aligns the cost profile with embeddings APIs, favoring high-frequency infrastructure automation over generative chat. Enterprises routing customer service tickets, classifying inbound documents, or filtering user-uploaded imagery no longer bear generation-side token penalties.
| API Endpoint | Output Format | Primary Use Case | Pricing Model |
|---|---|---|---|
/v1/decisions |
Typed answers (Predicate, Choice, Score) | Classification, routing, prioritization | $0.10 per 1M input tokens; no output charges |
| Responses API | Free-form text or structured JSON schemas | Content generation, extraction, function calling | Standard input and output token billing |
Implementation Mechanics and Endpoint Architecture
Requests to the POST /v1/decisions endpoint require three core components: the model parameter set to gpt-6-luna, a shared input field containing text or inline base64 data URLs, and a questions array defining the evaluation criteria. Hosted HTTP or HTTPS image URLs and file IDs are not supported by the endpoint.
Developers must assign a unique name to each question, which the API echoes back in the response array to simplify mapping. For instance, evaluating a customer complaint for department routing utilizes a choice structure with explicitly defined descriptions for billing, technical support, and shipping. If an input falls outside predefined categories, developers can include a fallback option like “other” to direct unhandled items to a general review queue.
From DevDay Preview to Public Beta Rollout
The Decisions API first surfaced on September 29, 2026, during OpenAI DevDay as a limited-invitation preview. Winzheng.com reported that the feature occupied less than a minute of the keynote stage, appearing alongside auxiliary tools such as Dots continuously running agents, the GPT-6.1 Sol reasoning model, and Codex Security Cloud. Exactly one week later, the tool transitioned from invitation-only status to an open public beta for all developers.
OpenAI expects the service to achieve General Availability in the coming weeks. While gpt-6-luna remains the sole supported model during the beta phase, the structured evaluation framework establishes a permanent paradigm for deterministic machine learning operations where binary and categorical decisions supersede prose generation.
- Silver Pines receives positive reviews ahead of October 8 release
- Repurpose old devices for tech hobbies like retro gaming
- Turning Garlic Peels Into Power: A New Self-Powered Home Security Sensor (world-today-journal.com)
- Rising Premiums Prompt Discussion of Universal Public Health Insurance Model (world-today-news.com)