
The entry-level Luna is now heavily discounted, with the input price per million tokens dropping from $1 to $0.20 and the output price falling from $6 to $1.20.
In RMB, the input price decreased from approximately ¥6.8 to ¥1.35, and the output price decreased from approximately ¥40.6 to ¥8.1.
Terra, positioned as the primary tool for daily tasks, has also dropped by 20%, with input and output prices reduced to $2 and $12, respectively.
Converted to RMB, they are approximately ¥13.5 and ¥81.2, respectively.
Only the flagship model, Sol, maintains its original price, but now offers a Fast mode that reaches up to 2.5 times the speed of the standard mode—without any additional cost.

After Codex drastically reset its usage limits, they immediately followed up with a price cut—OpenAI, who are you trying to outcompete?
This price war has directly intensified.
OpenAI stated that the price reduction is because GPT-5.6 Sol saved money on its own.
GPT-5.6 Sol participates in its own improvement, achieving a leap in efficiency—here’s a special thank-you to our new and existing users~
You're also into this whole "step on your left foot to levitate" routine, OpenAI.

.
Two reduced in price, one accelerated
Currently, the API prices for the three GPT-5.6 models are:
GPT-5.6 Luna: $0.20 per million tokens for input, $1.20 per million tokens for output;
GPT-5.6 Terra: $2 per million tokens input, $12 per million tokens output;
GPT-5.6 Sol: $5 per million tokens for input, $30 per million tokens for output.

After the price drop, Luna experienced a sharp plunge.
Its input price has reached parity with the previous generation's GPT-5.4 Nano, and its output price is even $0.05 cheaper.
However, the two models do not perform exactly the same tasks; GPT-5.4 nano is primarily designed for simple, high-frequency tasks such as classification, information extraction, and ranking.
GPT-5.6 Luna is positioned by OpenAI as a model designed for cost-sensitive workloads, capable of calling tools, handling long contexts, and executing multi-step workflows.
In simple terms, OpenAI is pricing the models that power agents at the same price point as the simple small models of the past.
Terra remained relatively restrained, dropping 20%.
Sol remains at the original price, with a new Fast mode added that reaches up to 2.5 times the speed of the standard mode, at twice the price of the standard mode, while maintaining the same model intelligence level.
It will also replace the previous Priority Processing. API requests that previously used the priority tag will automatically switch to Fast mode.
The official statement indicates that this adjustment will also apply to Codex and ChatGPT Work.
The subscription price will not be reduced, and the total quota for users will not increase, but the quota consumed when calling Terra and Luna will be correspondingly reduced.
The same subscription fee allows the model to perform more tasks.
OpenAI also replaced the Auto-review model in ChatGPT and Codex CLI from GPT-5.4 to GPT-5.6 Luna.
Auto-review is responsible for automatically checking code and modification results in the background, making it a typical high-frequency agent task.
Combined with model upgrades and the decline in Luna prices, OpenAI expects its costs to drop to approximately one-tenth of the original amount.
Affordable models are no longer limited to traditional tasks like classification and extraction; they are now entering areas such as code review and backend monitoring, which require judgment and tool invocation.
Currently, the positioning of the three models is:
Assign budget-sensitive, high-volume tasks to Luna;
Routine daily tasks are handled by Terra;
Continue using SOL for high-difficulty tasks;
If waiting feels too slow, double your money to buy faster speed.
GPT-5.6 begins participating in reducing its own costs.
Why the sudden price drop?
OpenAI stated that the reason is GPT-5.6 Sol participated in optimizing its own production system.
The efficiency it improves has been translated into discounted figures on the price list.
Specifically, GPT-5.6 Sol rewrote and optimized certain production-grade GPU kernels, designed and conducted hundreds of token generation experiments, and participated in training monitoring, stepping in when issues arose.
The final results were also outstanding: optimizing GPU kernels reduced the end-to-end service cost of GPT-5.6 by 20%; improving speculative decoding increased token generation efficiency by over 15%.

A GPU kernel can be understood as the low-level program that executes specific computational tasks on the GPU.
The model itself doesn't need to change; if these programs are written more efficiently, the same GPU can complete computations faster, handle more requests, and reduce the cost per invocation.
Speculative decoding generates potential next tokens in batches during answer creation, then submits them to the main model for verification. The more accurate the guesses, the fewer sequential generation and waiting steps are required.
Both optimizations maintain the model's capabilities while enabling the same model to run faster and at a lower cost.
However, this is not yet GPT-5.6 fully upgrading itself autonomously.
OpenAI specifically emphasized that the entire process is still human-led.
Models can write code, run experiments, and monitor anomalies, but determining goals, assessing whether results are usable, and deciding whether code enters production still require a "Human in the Loop."
OMT
Finally, there is one more new change.
This price reduction has a specific focus: Agent workflows.
The Luna with the deepest discounts isn't just a traditional small model used for classification and extraction.
According to OpenAI's positioning, it can invoke tools and execute multi-step tasks, primarily targeting high-frequency, cost-sensitive workloads.
Therefore, the 80% price drop in Luna also lowers the barrier to long-term operation for Agents.
Enable review, verification, and monitoring tasks that were previously too costly to perform frequently to become standard steps in your workflow.
Combined with GPT-5.6 participating in optimizing its own production system, a new cycle emerges:
Model participation improves operational efficiency and reduces costs; after prices drop, the model is adopted in more high-frequency workflows; increased usage scales drive the next round of efficiency optimizations.
The pricing flywheel from OpenAI is already taking shape.
After all these actions, the pressure is now squarely on Company A:
It's your turn.

Reference links: [1] https://x.com/OpenAI/status/2082878156483219672 [2] https://x.com/sama/status/2082880720989532597
This article is from the WeChat public account "Quantum Bit," authored by: Focused on Frontier Technology
