Price cuts on smaller and mid-tier models

OpenAI announced on Thursday that it is reducing API pricing for two models in its GPT-5.6 family, effective July 30. The lightweight GPT-5.6 Luna model dropped 80% to US$0.20 per million input tokens and US$1.20 per million output tokens, while the mid-tier GPT-5.6 Terra fell 20% to US$2 per million input tokens and US$12 per million output tokens, according to the company's blog post as reported by multiple outlets. CEO Sam Altman disclosed the cuts in a post on X. Prior to the change, Terra cost US$2.50 per million input and US$15 per million output, and Luna cost US$1 and US$6 respectively.
Flagship Sol model and new Fast mode
Pricing for OpenAI's flagship GPT-5.6 Sol was left unchanged. Instead, the company introduced a new "Fast mode" for Sol that replaces the previous Priority Processing API option, which it said boosts performance speed up to 2.5 times. Existing API requests tagged "priority" will continue to work, according to OpenAI's blog post and follow-on coverage.
Efficiency gains underpin the move
OpenAI attributed the lower prices to improvements in serving efficiency across its training and inference stack, including software and GPU-infrastructure optimizations. The company said GPT-5.6 Sol helps optimize the production GPU kernels used to run AI workloads, reducing inference costs without compromising model performance. The trio of GPT-5.6 models was introduced earlier this month.
Competitive pressure from Chinese and US rivals
The cuts come as US AI labs face mounting competition from cheaper Chinese alternatives. Reuters reporting, citing analysts, identified open-source Chinese rival Z.ai's GLM-5.2 as a model that "nearly match[es] their performance at a lower cost." The South China Morning Post, citing rankings from research firm Artificial Analysis, said GPT-5.6 Luna moved into the "most attractive" tier on intelligence-per-dollar metrics, placing it above rivals the outlet identified as Zhipu AI's GLM-5.2 and MiniMax's M3; Reuters attributed the leading Chinese competitor solely to Z.ai, reflecting a discrepancy between the two reports. On the US side, OpenAI's new pricing undercuts Anthropic's mid-tier Claude Sonnet 4.6, which lists at US$3 per million input tokens and US$15 per million output tokens.
Enterprise spending implications
Analysts said the price reductions are more likely to accelerate AI deployment than to reduce overall enterprise spending. Lower per-token pricing makes it easier to move pilots into production, expanding the volume of AI work enterprises can perform for a given budget. OpenAI also reduced the number of usage credits that Luna and Terra consume within ChatGPT Work and Codex, effectively raising the amount of work enterprise subscribers can do without paying more. Industry observers expect inference costs across providers, including Anthropic, Google and Microsoft, to keep declining over the next 24 months. Analysts added that broader usage-based pricing is leaving companies with unpredictable bills even as per-token costs fall.
Next milestone
Analysts expect the pricing pressure on US labs to continue, with attention turning to whether Anthropic, Google and Microsoft respond with their own adjustments and how OpenAI's smaller-model margins hold up ahead of any initial public offering activity.
Share this article







