OpenAI Cuts GPT-5.6 Prices by Up to 80% to Rival Google and Anthropic

OpenAI has slashed the API price of its GPT-5.6 Luna model by 80% to $1.40 per million total tokens, igniting an enterprise inference cost war alongside rival releases from Anthropic and Google.

The Economics of the Frontier Inference Drop

Begun, the AI price wars have. OpenAI’s structural recalibration of its GPT-5.6 lineup drops Luna—the smallest and fastest model in the architecture—from a combined $7 per million tokens down to $0.20 per million input tokens and $1.20 per million output tokens. This aggressive down-market push arrives just days after Anthropic rolled out its performant Claude Opus 5 at the same price point as its predecessor, and Google introduced its low-latency Gemini 3.6 Flash and Gemini 3.5 Flash-Lite variants.

OpenAI co-founder and CEO Sam Altman announced the changes on X as “major price cuts today.” The adjustments target the exact operational metric enterprise buyers care about: the total cost of executing production workloads rather than raw token sticker prices.

Luna now undercuts Google’s Gemini 3.5 Flash-Lite, which sits at a combined $2.80 per million input and output tokens, and positions itself below Google’s Gemini 3.6 Flash at $9 per million tokens. While pure-play low-cost providers like Xiaomi with MiMo-V2.5 Flash and DeepSeek remain cheaper on a raw token basis, OpenAI has successfully dragged its proprietary frontier-series architecture into the low-cost tier.

Where the GPT-5.6 Tiers Stand Now

Beyond the headline-grabbing Luna reduction, OpenAI also trimmed the mid-tier GPT-5.6 Terra model by 20%. Terra’s combined price dropped from $17.50 to $14 per million tokens, aligning it directly with Google’s Gemini 3.1 Pro Preview pricing for context windows under 200,000 tokens.

OpenAI Slashes GPT-5.6 Luna Prices Dramatically

As noted on X by Krea AI’s Nic Dunz, Terra offers the same intelligence as OpenAI’s older GPT-5.4 model for about 1/13th the cost, while maintaining a wide performance gap over lower-tier alternatives. Meanwhile, the flagship GPT-5.6 Sol model remains anchored at $5 per million input tokens and $30 per million output tokens for its Standard mode. However, OpenAI introduced a premium Sol Fast mode at twice that rate—$10 per million input tokens and $60 per million output tokens—which delivers up to 2.5 times the throughput without degrading underlying model intelligence.

Model Tier Input ($/1M Tokens) Output ($/1M Tokens) Combined Total ($/1M Tokens)
GPT-5.6 Luna $0.20 $1.20 $1.40
GPT-5.6 Terra $2.00 $12.00 $14.00
GPT-5.6 Sol — Standard $5.00 $30.00 $35.00
GPT-5.6 Sol — Fast $10.00 $60.00 $70.00

Performance Parity and the Pareto Curve

Third-party evaluations from Artificial Analysis indicate that OpenAI’s models outperform competing Google variants, with Luna scoring higher on intelligence metrics than Gemini 3.6 Flash and the older Gemini 3.1 Pro.

AI coding startup Cognition highlighted on X that the GPT-5.6 series now sits squarely “on the pareto curve of price/performance efficiency.” Their published analysis demonstrated the model lineup shifting left on a coordinate plane of intelligence versus cost, proving that top-tier reasoning capabilities are no longer tethered exclusively to maximum API pricing.

At the high end, Anthropic chose a different strategic lever with Claude Opus 5. Maintaining an unchanged sticker price of $5 per million input tokens and $25 per million output tokens—matching the rates of Opus 4.8—Anthropic delivered near-Fable 5 performance at half the operating cost of that larger system. By integrating adjustable effort settings, Anthropic allows developers to dynamically scale reasoning depth against token expenditure.

The Shift Toward Production Economics

Enterprise platform architects are no longer evaluating models solely on raw benchmark scores. High-volume deployments—ranging from automated coding agents to real-time document classification and customer support routing—magnify minor discrepancies in inference pricing.

Google has similarly oriented its Gemini 3.6 Flash and Flash-Lite releases around agent deployment economics, emphasizing reductions in output token consumption and tool-call overhead. With Gemini 3.6 Flash reportedly utilizing 17% fewer output tokens than its predecessor on benchmark indexes, and savings climbing to 65% on extended engineering tasks, the vector of competition has clearly pivoted from sheer capability access to cost predictability.

Sol retains the crown for complex reasoning and agentic workflows, Terra handles general production tasks, and Luna now stands ready to absorb high-frequency, low-latency API calls across the ecosystem.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

Kent Allrounder to Leave After Four Years; Warwickshire Sign Championship Replacements

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.