Why OpenAI’s 80% Price Cut Could Trigger A Race To The Bottom In AI
OpenAI has cut the price of GPT-5.6 Luna by 80%, reducing its API cost from $1 to 20 cents per million input tokens and from $6 to $1.20 per million output tokens. GPT‑5.6 Terra will cost 20% less while the most recent GPT-5.6 Sol is unchanged. OpenAI says improvements in system efficiency made the reductions possible.
The cuts arrived as ChatGPT reached an extraordinary scale. Sensor Tower estimated in June that the ChatGPT app had surpassed one billion monthly active users, making it the fastest consumer application to reach that level. OpenAI has subsequently reported more than one billion active users across its services and two million business customers. These figures measure accounts or application activity rather than one billion paying customers, yet they still mark a major transition.
The price reduction reveals three larger changes.
Competition from open-weight models is pushing proprietary providers toward lower prices.
AI adoption is moving considerably faster than earlier computing revolutions, leaving institutions little time to adjust.
The foundation-model market is also beginning to resemble an infrastructure industry in which scale, capital and operating efficiency favor a limited number of large providers.
Challenge from Open-Weight Models
OpenAI’s decision reflects growing pressure from enterprises that are scrutinizing the cost of AI deployment. A model can appear inexpensive when tested through a chatbot, yet become costly when embedded in a customer-service system, coding platform or research workflow that processes billions of tokens. AI Agents increase that expense further because they may search, reason, call tools and revise their work repeatedly before completing one task. At that scale, the difference between $1 and 20 cents per million input tokens becomes a serious procurement issue.
Open-weight models have strengthened the buyer’s bargaining position. Moonshot AI’s Kimi K3, DeepSeek V4 and Z.ai’s GLM-5.2 demonstrate that capable models can be distributed at low prices and, in many cases, downloaded or adapted by outside developers. Kimi K3, released in July, is a mixture-of-experts model with 2.8 trillion total parameters, 104 billion activated parameters and a context window of one million tokens. Its developers also released the model weights, allowing companies and researchers to inspect, modify and deploy it independently.
Their importance extends beyond benchmark rankings. Open weights give developers additional choices over hosting, customization,........
