OpenAI Launches Decisions API in Public Beta Powered by GPT-6 Luna

The OpenAI Decisions API entered public beta on October 6, 2026, powered exclusively by the gpt-6-luna model and accessible via the dedicated POST /v1/decisions endpoint, running up to 10 times faster than the Responses API by returning typed answers directly without token-by-token text generation.

Why GPT-6 Luna Skips Token Generation

Large language models traditionally burn computational cycles generating text word by word, even when an application merely needs a binary judgment or a categorical routing choice. developers.openai.com reported that the Decisions API cuts out this unnecessary generation step entirely. Instead of forcing a model to write out a reasoning chain and a prose response, the endpoint evaluates text, images, or both, and immediately returns structured, typed data.

The performance gain comes from eliminating the generation bottleneck, not merely scaling underlying hardware compute. For high-throughput infrastructure processing tens of thousands of requests daily, removing output token generation changes the fundamental economics of automated classification.

The API handles requests through three distinct question types. Predicate questions evaluate conditions—such as checking a product photo for visible damage—and return a probability estimation between 0 and 1. Choice questions select a single option from an unordered list, returning probabilities and a confidence score for categories like billing, technical support, or account access. Score questions rate inputs against ordered levels, calculating a probability-weighted average of numeric level indices that can fall between discrete steps, such as severity ratings.

OpenAI Decisions API Opens Public Beta: Powered by GPT-6 Luna, Up to 10x Faster Decisions, Input-Only Pricing
Photo: winzheng.com

Pricing Structures and Enterprise Token Costs

The financial model of the Decisions API mirrors infrastructure utilities rather than conversational chat interfaces. Input tokens cost $0.10 per million tokens. Developers pay exclusively for input tokens, with zero cache-read, cache-write, or output-token charges. Regional processing premiums and long-context input pricing multipliers apply for eligible deployments.

By abandoning output-token billing for this endpoint, OpenAI aligns the cost profile with embeddings APIs, favoring high-frequency infrastructure automation over generative chat. Enterprises routing customer service tickets, classifying inbound documents, or filtering user-uploaded imagery no longer bear generation-side token penalties.

API Endpoint Output Format Primary Use Case Pricing Model
/v1/decisions Typed answers (Predicate, Choice, Score) Classification, routing, prioritization $0.10 per 1M input tokens; no output charges
Responses API Free-form text or structured JSON schemas Content generation, extraction, function calling Standard input and output token billing

Implementation Mechanics and Endpoint Architecture

Requests to the POST /v1/decisions endpoint require three core components: the model parameter set to gpt-6-luna, a shared input field containing text or inline base64 data URLs, and a questions array defining the evaluation criteria. Hosted HTTP or HTTPS image URLs and file IDs are not supported by the endpoint.

Developers must assign a unique name to each question, which the API echoes back in the response array to simplify mapping. For instance, evaluating a customer complaint for department routing utilizes a choice structure with explicitly defined descriptions for billing, technical support, and shipping. If an input falls outside predefined categories, developers can include a fallback option like “other” to direct unhandled items to a general review queue.

From DevDay Preview to Public Beta Rollout

The Decisions API first surfaced on September 29, 2026, during OpenAI DevDay as a limited-invitation preview. Winzheng.com reported that the feature occupied less than a minute of the keynote stage, appearing alongside auxiliary tools such as Dots continuously running agents, the GPT-6.1 Sol reasoning model, and Codex Security Cloud. Exactly one week later, the tool transitioned from invitation-only status to an open public beta for all developers.

OpenAI expects the service to achieve General Availability in the coming weeks. While gpt-6-luna remains the sole supported model during the beta phase, the structured evaluation framework establishes a permanent paradigm for deterministic machine learning operations where binary and categorical decisions supersede prose generation.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

CamelBak Podium Stainless Insulated Bike Bottle is 36% Off for Prime Day

Jack White and California officials slam Trump over Iran rally remarks