OpenAI cuts GPT-5.6 Terra prices 20% and Luna 80%

OpenAI cuts GPT-5.6 Terra prices 20% and Luna 80%

OpenAI lowered GPT-5.6 Terra and Luna API prices, introduced a premium Fast Mode for Sol, and highlighted efficiency gains as developers weigh costs against lower-priced Chinese open-weight rivals.

SOL

Fact Check
The official OpenAI announcement (Advancing the price-performance frontier with GPT-5.6) directly confirms the 20% Terra cut and 80% Luna cut, and explicitly states that while subscription prices and quota budgets for ChatGPT and Codex remain the same, Terra and Luna usage now consumes fewer credits — supporting the claim that the cuts flow through to subscription quota calculations. OpenAI's official X account, CNBC, and Quartz/Yahoo all corroborate identical percentages and per-token pricing. The only minor nuance: the title's phrase 'ChatGPT Work' is not an exact product name in the sources, which reference 'ChatGPT' and 'Codex'; this does not materially affect the core claim.
Summary

OpenAI cut GPT-5.6 Luna prices by 80% and GPT-5.6 Terra by 20% on July 30, roughly three weeks after the lineup reached users, while keeping Sol's standard pricing unchanged and replacing Priority Processing with a new Fast Mode. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, Terra costs $2 and $12, and Sol remains at $5 and $30, with Fast Mode priced at $10 and $60 per million tokens and advertised as running Sol up to 2.5 times faster than standard processing. OpenAI said existing Priority Processing users will be migrated automatically and that Terra and Luna will consume fewer credits in ChatGPT Work and Codex without changes to subscription prices or quota budgets. The company attributed some savings to work done by Sol under human supervision, saying the model rewrote production kernels, ran token-generation experiments and monitored training runs, lifting token-generation efficiency by more than 15% and cutting serving costs by 20%. The pricing move comes as lower-cost Chinese open-weight models from DeepSeek, Zhipu and Moonshot continue to attract developers handling large AI workloads, and after concerns over operating costs and model behavior sharpened scrutiny of enterprise AI deployments.

Terms & Concepts
  • input tokens: Units of text sent into an AI model as prompts or other data for processing and billing.
  • Fast Mode: OpenAI's higher-priced API processing option that is designed to deliver quicker model responses than standard processing.
  • open-weight models: AI models released with downloadable parameters, allowing developers to run them outside a provider's hosted API.