História
agosto 5, 2026
OpenAI’s 80% Luna Cut Turns AI’s Price War Into an ROI Test
OpenAI says lower GPT-5.6 prices come from real efficiency gains across its stack. But as agentic workloads inflate token use and cheaper rivals press in, enterprise buyers are likely to focus less on rates than on finished work.
OpenAI has cut the price of its cheapest GPT-5.6 model by 80%, betting that a dramatic reduction in API costs will make its case for “abundant intelligence.” The harder question is whether cheaper tokens will finally translate into better returns for companies buying AI.
The move came only about three weeks after GPT-5.6 launched. On Thursday, OpenAI lowered Luna to $0.20 per million input tokens and $1.20 per million output tokens, while cutting mid-tier Terra by 20%; flagship Sol was spared a discount, though it gained a faster API option.1 Greg Brockman framed Luna as “by far the most price-efficient model in its class,” presenting the cut as the product of research into efficient intelligence rather than a concession in a bruising market.
2
OpenAI’s broader explanation is that pricing is no longer the useful unit of comparison. In its account, customers are buying completed tasks—not raw token volume—and the relevant measure is the cost of a successful outcome after retries, human oversight and errors.3 The company says improvements to routing, context management and production software helped reduce serving costs by 20%, while speculative decoding lifted token-generation efficiency by more than 15%.4
That is the optimistic reading: lower costs expand the work AI can economically perform, feeding adoption and funding further gains. OpenAI says its models now reach more than one billion active users and two million businesses, scale it argues helps that cycle compound.3
The market reading is less charitable. Cheaper Chinese open-weight models have raised pressure on closed-model vendors to justify their premiums, while long-running agentic tasks consume far more tokens and have made customers increasingly price-sensitive.1 As analyst Jacob Bourne put it, “the era of tokenmaxxing is over”: enterprises can burn through tokens without getting value back.5
Altman’s answer was characteristically bullish—“we want to offer the best price/intelligence tradeoff at every level.”
6 The next test is whether OpenAI can prove that tradeoff in business outcomes, not merely on a price sheet.