OpenAI Reduces GPT-5.6 Luna and Terra API Prices, Introduces Faster Sol Mode

icon MarsBit
Share
AI summary iconSummary
OpenAI reduces GPT-5.6 Luna and Terra API costs, with Luna’s input and output fees each down 80% and Terra’s reduced by 20%. The Sol model now offers a 2.5x faster Fast mode at the same price. Efficiency improvements from Sol—including speculative decoding and GPU kernel optimizations—help lower costs. As the Fear and Greed Index indicates rising market confidence, these reductions may increase adoption of AI agents and influence crypto price dynamics in high-frequency trading.

Agent

The entry-level Luna is now heavily discounted, with the input price per million tokens dropping from $1 to $0.20 and the output price falling from $6 to $1.20.

In RMB, the input price decreased from approximately ¥6.8 to ¥1.35, and the output price decreased from approximately ¥40.6 to ¥8.1.

Terra, positioned as the primary tool for daily tasks, has also dropped by 20%, with input and output prices reduced to $2 and $12, respectively.

Converted to RMB, they are approximately ¥13.5 and ¥81.2, respectively.

Only the flagship model, Sol, maintains its original price, but now offers a Fast mode that reaches up to 2.5 times the speed of the standard mode—without any additional cost.

Agent

After Codex drastically reset its usage limits, they immediately followed up with a price cut—OpenAI, who are you trying to outcompete?

This price war has directly intensified.

OpenAI stated that the price reduction is because GPT-5.6 Sol saved money on its own.

GPT-5.6 Sol participates in its own improvement, achieving a leap in efficiency—here’s a special thank-you to our new and existing users~

You're also into this whole "step on your left foot to levitate" routine, OpenAI.

Agent

.

Two reduced in price, one accelerated

Currently, the API prices for the three GPT-5.6 models are:

GPT-5.6 Luna: $0.20 per million tokens for input, $1.20 per million tokens for output;

GPT-5.6 Terra: $2 per million tokens input, $12 per million tokens output;

GPT-5.6 Sol: $5 per million tokens for input, $30 per million tokens for output.

Agent

After the price drop, Luna experienced a sharp plunge.

Its input price has reached parity with the previous generation's GPT-5.4 Nano, and its output price is even $0.05 cheaper.

However, the two models do not perform exactly the same tasks; GPT-5.4 nano is primarily designed for simple, high-frequency tasks such as classification, information extraction, and ranking.

GPT-5.6 Luna is positioned by OpenAI as a model designed for cost-sensitive workloads, capable of calling tools, handling long contexts, and executing multi-step workflows.

In simple terms, OpenAI is pricing the models that power agents at the same price point as the simple small models of the past.

Terra remained relatively restrained, dropping 20%.

Sol remains at the original price, with a new Fast mode added that reaches up to 2.5 times the speed of the standard mode, at twice the price of the standard mode, while maintaining the same model intelligence level.

It will also replace the previous Priority Processing. API requests that previously used the priority tag will automatically switch to Fast mode.

The official statement indicates that this adjustment will also apply to Codex and ChatGPT Work.

The subscription price will not be reduced, and the total quota for users will not increase, but the quota consumed when calling Terra and Luna will be correspondingly reduced.

The same subscription fee allows the model to perform more tasks.

OpenAI also replaced the Auto-review model in ChatGPT and Codex CLI from GPT-5.4 to GPT-5.6 Luna.

Auto-review is responsible for automatically checking code and modification results in the background, making it a typical high-frequency agent task.

Combined with model upgrades and the decline in Luna prices, OpenAI expects its costs to drop to approximately one-tenth of the original amount.

Affordable models are no longer limited to traditional tasks like classification and extraction; they are now entering areas such as code review and backend monitoring, which require judgment and tool invocation.

Currently, the positioning of the three models is:

Assign budget-sensitive, high-volume tasks to Luna;

Routine daily tasks are handled by Terra;

Continue using SOL for high-difficulty tasks;

If waiting feels too slow, double your money to buy faster speed.

GPT-5.6 begins participating in reducing its own costs.

Why the sudden price drop?

OpenAI stated that the reason is GPT-5.6 Sol participated in optimizing its own production system.

The efficiency it improves has been translated into discounted figures on the price list.

Specifically, GPT-5.6 Sol rewrote and optimized certain production-grade GPU kernels, designed and conducted hundreds of token generation experiments, and participated in training monitoring, stepping in when issues arose.

The final results were also outstanding: optimizing GPU kernels reduced the end-to-end service cost of GPT-5.6 by 20%; improving speculative decoding increased token generation efficiency by over 15%.

Agent

A GPU kernel can be understood as the low-level program that executes specific computational tasks on the GPU.

The model itself doesn't need to change; if these programs are written more efficiently, the same GPU can complete computations faster, handle more requests, and reduce the cost per invocation.

Speculative decoding generates potential next tokens in batches during answer creation, then submits them to the main model for verification. The more accurate the guesses, the fewer sequential generation and waiting steps are required.

Both optimizations maintain the model's capabilities while enabling the same model to run faster and at a lower cost.

However, this is not yet GPT-5.6 fully upgrading itself autonomously.

OpenAI specifically emphasized that the entire process is still human-led.

Models can write code, run experiments, and monitor anomalies, but determining goals, assessing whether results are usable, and deciding whether code enters production still require a "Human in the Loop."

OMT

Finally, there is one more new change.

This price reduction has a specific focus: Agent workflows.

The Luna with the deepest discounts isn't just a traditional small model used for classification and extraction.

According to OpenAI's positioning, it can invoke tools and execute multi-step tasks, primarily targeting high-frequency, cost-sensitive workloads.

Therefore, the 80% price drop in Luna also lowers the barrier to long-term operation for Agents.

Enable review, verification, and monitoring tasks that were previously too costly to perform frequently to become standard steps in your workflow.

Combined with GPT-5.6 participating in optimizing its own production system, a new cycle emerges:

Model participation improves operational efficiency and reduces costs; after prices drop, the model is adopted in more high-frequency workflows; increased usage scales drive the next round of efficiency optimizations.

The pricing flywheel from OpenAI is already taking shape.

After all these actions, the pressure is now squarely on Company A:

It's your turn.

Agent

Reference links: [1] https://x.com/OpenAI/status/2082878156483219672 [2] https://x.com/sama/status/2082880720989532597

This article is from the WeChat public account "Quantum Bit," authored by: Focused on Frontier Technology

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.