Skip to content
LIVE
The Executives BriefThe Executives BriefBeta

OpenAI slashes GPT-5.6 Luna prices 80% to $1.40 per million tokens

The fastest GPT-5.6 tier drops into the low-cost tier, while Terra gets a smaller cut and Sol adds a premium Fast mode.

ByOmar Al-BalawiTechnology Correspondent, The Executives Brief
·4 min read
OpenAI slashes GPT-5.6 Luna prices 80% to $1.40 per million tokens
Executive summary

OpenAI is cutting prices across its GPT-5.6 frontier model lineup, slashing GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, and adding a new Fast mode for GPT-5.6 Sol. The move shifts OpenAI's frontier positioning toward cost-per-request competition, right after rival launches that emphasized cheaper agent deployment.

OpenAI just dropped the price on GPT-5.6 Luna by 80%. The company says the smallest and fastest model in its GPT-5.6 “frontier” family now costs $0.20 per million input tokens and $1.20 per million output tokens, for a combined $1.40 per million tokens. That is a massive shift from its prior $7 per million combined input-plus-output price, dragging Luna much closer to the market’s low-cost commercial tier.

This is not happening in a vacuum. OpenAI made these changes only days after Anthropic released Claude Opus 5 at the same price as Opus 4.8, and after Google introduced Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, both engineered around lower inference costs, faster execution, and more efficient agent workloads. In other words, OpenAI is responding to a pricing narrative that rivals are already winning: lower token usage and fewer “expensive steps” for long-running work.

So what exactly changed in the GPT-5.6 lineup? Alongside Luna’s 80% cut, OpenAI also reduced GPT-5.6 Terra, the mid-tier model, by 20%. Terra now costs $2 per million input tokens and $12 per million output tokens, for a combined $14 per million tokens. OpenAI did not cut its flagship “Sol” baseline price: Sol Standard remains at $5 per million input tokens and $30 per million output tokens, for $35 per million tokens total. But OpenAI is adding Sol Fast mode at twice the Standard price, $10 per million input tokens and $60 per million output tokens. OpenAI says Sol Fast delivers up to 2.5 times the throughput without changing the model’s underlying intelligence.

For operators and finance people, the interesting part is not only the sticker price. It is what the pricing implies about how OpenAI expects customers to allocate workloads across tiers. OpenAI describes GPT-5.6 Sol as aimed at complex reasoning-heavy and agentic tasks, including advanced coding, multi-step planning, and tool-using systems. Terra is positioned for general production use where teams want a balance of capability and efficiency. Luna is positioned for high-throughput, low-latency tasks like summarization, classification, routing, and lightweight real-time assistants where cost per request is the primary constraint. With Luna’s combined price now at $1.40 per million tokens, OpenAI is effectively lowering the bar for teams that previously had to choose between “frontier” quality and “low-cost inference” economics.

The Terra cut also matters because it narrows how wide the internal tier gaps can feel. Terra’s combined price moves from $17.50 down to $14 per million tokens, which at that level matches Google’s Gemini 3.1 Pro Preview pricing for context windows of 200,000 tokens or less. And it undercuts OpenAI’s own GPT-5.4 and places Terra into clearer competition with other “production” tiers. Meanwhile, the Sol Fast move goes the other direction: instead of making Sol cheaper, OpenAI is making latency-sensitive throughput a paid premium. That creates a cleaner segmentation: cheaper tokens for throughput-light tasks at the Luna tier, balanced pricing for general production at Terra, and a “pay more for speed” dial at Sol Fast.

The broader market pressure is visible in the sequencing. Google introduced Gemini 3.6 Flash priced at $1.50 per million input tokens and $7.50 per million output tokens, and Gemini 3.5 Flash-Lite priced at $0.30 per million input tokens and $2.50 per million output tokens. Google framed these models around agent economics, arguing that lower token usage, fewer reasoning steps, and reduced tool calls could lower total cost for long-running engineering and knowledge-work tasks. The reporting also notes Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching as high as 65% on some long-horizon engineering workloads. OpenAI’s Luna cut is designed to win a similar cost-per-usage argument, but with its own claim of cost-per-intelligence advantage. Third-party analysis outfits such as Artificial Analysis are cited as showing OpenAI’s models as more performant than Google’s, including Luna outperforming Gemini 3.6 Flash and older Gemini 3.1 Pro, making “cost-per-intelligence” more favorable to OpenAI.

Then there is Anthropic. Claude Opus 5 is described as remaining about as performant as GPT-5.6 Sol, but 6% cheaper than OpenAI on the reported pricing comparison. Anthropic says Opus 5 costs $5 per million input tokens and $25 per million output tokens, the same as Opus 4.8, and Anthropic says it delivers nearly all the intelligence of its more expensive Fable 5 at roughly half the cost. Unlike OpenAI’s Luna and Terra changes, Anthropic did not reduce the Opus API sticker price. Instead, the narrative is that Anthropic effectively lowered price per unit of capability by moving from Opus 4.8 to a more capable version.

For executives, boards, and finance leads, this pricing wave is a signal: frontier model families are no longer just competing on intelligence. They are competing on how cheaply customers can deploy useful systems at scale, including agents that burn tokens across long workflows. When OpenAI pushes Luna into the low-cost tier and reserves Sol speed as a premium add-on, it is forcing competitors to defend not only performance, but also allocation strategy. If you are building products, forecasting unit economics, or negotiating enterprise deals, these shifts change what “good” looks like for your marginal request. In a world where token economics can decide whether an app ships or stalls, OpenAI is moving fast enough to make pricing itself a battlefield.

Executive ActionsLocked

This story's Key Insights and Take-aways are locked.

Create a free account to unlock Executive Actions for one credit.

Register to Unlock

Always free for Executives Club members. Join the Club

More in Technology