
If anyone is still spending the most money to invoke OpenAI's most powerful model today.
OpenAI would instead suggest he switch to another.
On July 30, OpenAI issued a price adjustment announcement.
The GPT-5.6Luna model is now 80% off, and the Terra model is 20% off.
Seeing this message, it's easy to focus on the price war that has begun in Silicon Valley.
However, after carefully reviewing the official technical documentation and API usage guide, you'll realize that the truly significant actions lie not in the price numbers themselves.
This is essentially the first time OpenAI has told users that many tasks don't require the most powerful models.
The official provided a specific recommendation: for complex tasks, first use GPT-5.6Sol to complete requirements analysis and solution design, then assign execution, coding, and testing to Luna.
The most expensive model handles thinking, while the cheapest model does the work. Two years ago, this strategy would have seemed equivalent to business suicide.
Keep in mind that, in the past, the entire AI industry was desperately telling the market that their models were the smartest.
Today, however, OpenAI stated that you don't always need to buy the most expensive option.
This matter is far more important than a price reduction.
I. The Unspoken Agreement Between Silicon Valley's Two Titans
First, let's look at what happened over the past two weeks.
On July 30, OpenAI adjusted its pricing: Luna decreased by 80%, Terra decreased by 20%, and Sol did not lower its price but introduced a Fast mode, which can be up to 2.5 times faster than the standard mode at double the price, while maintaining the same level of intelligence.
The top-tier models remain unchanged. The real volume growth is coming from mid- and low-end models.
One week earlier.
On July 24, Anthropic did something nearly identical: Claude Opus 5 was released, with input at $5 per million tokens and output at $25 per million tokens—exactly half the price of Fable 5.
Rather than highlighting performance breakthroughs, Anthropic emphasizes its cost-effectiveness: achieving reasoning capabilities nearly identical to Fable 5 at just half the price.
Just a month ago, Fable 5 was Anthropic’s flagship product, heavily promoted for its superior reasoning, longest context length, and highest price. Yet just one month later, Anthropic itself introduced a half-priced alternative as its new flagship.
If only one company did this, it could be seen as a product adjustment. But if two did it almost simultaneously, it’s not a coincidence.
They have all begun to proactively reduce the importance of their flagship models.
Second, the flagship handles promotion, while volume drives profits.
I’ve been wondering ever since why it’s happening now.
The answer is actually not complicated.
The greatest value of past flagship models wasn't generating profits, but demonstrating technological leadership. After GPT-4 was released, OpenAI’s valuation continued to rise. Each update to Claude allows Anthropic to redefine its technological position. Flagship models carry brand value.
But that's not where businesses actually spend their money.
A company makes millions of API calls daily. High-frequency tasks such as customer service, search, approval processes, code generation, and Agent execution account for the majority of token consumption. What enterprises care most about is not achieving the top benchmark, but rather the cost per task, stability, and return on investment.
When scaling to tens of millions of calls per day, the slight intelligence advantage of flagship models is instantly erased by the massive computational costs.
OpenAI's actions and statements mark a turning point: top-tier flagship models are no longer burdened with the responsibility of generating profits.
III. AI is entering the "mass-market" era
This scenario is something the automotive industry has already experienced.
Twenty years ago, the 7 Series defined BMW’s prestige, the S-Class established Mercedes-Benz’s luxury character, and the A8 set the foundation for Audi’s flagship image—flagship models determined a brand’s upper limit. Yet, the core models that consistently drove sales volume and generated profits have always been the BMW 3 Series, Mercedes-Benz C-Class, and Audi A4.
Later, it became even clearer. The Model S proved that Tesla could build cars; it was the Model 3 and Model Y that truly made Tesla a global automaker.
The flagship model is responsible for demonstrating capability, while the mass-market models handle scale.
AI has also begun entering this stage. Sol and Fable will continue to exist, focusing on pushing technological boundaries and setting new benchmarks. The entities increasingly taking on commercialization will become Luna, Terra, and Opus.
This time, OpenAI even publicly outlined the recommended process: Sol is responsible for planning, and Luna is responsible for execution.
This is no longer just about a single model—OpenAI is designing a system of specialized models. In the future, enterprises won’t buy just one model, but an entire ecosystem of models. What truly determines cost isn’t the chief architect, but the construction team working day in and day out.
Four, the model begins to optimize the model
There's another detail I find even more interesting than the price reduction itself.
OpenAI mentioned in its technical documentation that this price reduction is not solely due to purchasing more GPUs or because of increased scale. The real reason is that the model itself is now participating in optimizing the model.
Specifically, under the guidance of human engineers, Sol autonomously rewrote and optimized the underlying production kernel. It designed hundreds of experiments to improve token generation efficiency and even participated in monitoring the model training process, intervening directly when issues were detected.
As a result, end-to-end operational costs decreased by 20%, and Token generation efficiency improved by 15%.
This message is easy to overlook, but it carries significant meaning.
In the past, engineers improved efficiency. After the model went live, people once Click here Optimizing the inference framework, CUDA, caching strategies, and scheduling algorithms—achieving a 10+ percentage point improvement in efficiency over a year is already impressive.
Technology has now entered a new phase, with models taking over engineering optimizations for underlying code and computational resource scheduling, running and iterating continuously around the clock.
This is a self-reinforcing cycle that accelerates over time. The smarter the model, the more effectively it can participate in optimization. Faster optimization leads to quicker cost reductions. Lower costs drive higher usage volumes, which generate more data to further train the model.
I reviewed OpenAI’s price changes over the past two and a half years. GPT-4 initially launched at $30 per million tokens for input, GPT-4o dropped to $5, and GPT-4o mini went as low as $0.15. Today, Luna’s pricing aligns with the budget range of early mini models, yet its overall intelligence far surpasses the expensive GPT-4 from two years ago.
Over two years, the price has dropped by nearly two orders of magnitude. If the model continues to participate in optimizing itself, this curve will likely continue to decline.
What's truly terrifying isn't that the price dropped 80% today, but that costs have begun to develop self-sustaining downward momentum.
Five: From Who Is the Smartest to Who Is Most Worthy
I’m increasingly feeling that when people discuss AI competition today, they sometimes still rely on the framework of the previous stage.
For example, users often follow up by asking: Who is the smartest? GPT, Claude, Gemini, DeepSeek Every time a new model is released, the media immediately looks at the rankings—who’s number one wins.
However, recent moves by OpenAI and Anthropic have broken this pattern, sending a new signal to the market: the appeal of single-performance champions is waning.
OpenAI mentioned in its announcement a point I find particularly crucial: it recommends that developers match different models based on the importance of the task, the cost of errors, urgency, and scale.
Note that what is being discussed is no longer the model, but the task.
In the past, when a company deployed AI, it usually had only one option. Starting today, it’s more like building an organization: the most critical tasks go to Sol, routine operations are handled by Luna, and in the future, even lighter models may take on simpler tasks.
This is very similar to cloud computing back then. No one would store all their data on the most expensive SSDs. Hot data goes on SSDs, regular data on HDDs, and cold data on object storage. The discussion has never been about which hard drive is fastest—it’s always been about how to build the most cost-effective system overall.
DeepSeek Actually, it sensed this direction even earlier than Silicon Valley. Over the past six months, it has rarely emphasized being the smartest in the world, repeatedly focusing on just a few words: cheap, fast enough, and sufficient.
After a cache hit, the cost per million tokens drops to nearly negligible levels. It has consistently bet on one thing: businesses don’t need world champions—they need the deployment option with the highest overall ROI.
This is not a short-term price war; the recent moves by these two Silicon Valley giants confirm the inevitability of this business path.
Six: What OpenAI truly wants to sell is not the model
As of today, OpenAI’s strategy is clear: what it truly wants to sell is no longer the model itself, but the volume of API calls.
Previously, model manufacturers relied on high technological premiums to achieve high profit margins; now, the business model has shifted to gaining access to a massive user ecosystem through extremely low entry barriers.
Microsoft doesn't make its real profits because Windows is expensive—it's because every computer runs Windows. AWS doesn't profit because of high margins on individual servers—it's because countless applications run on it every day around the world.
The platform's revenue has always depended on penetration rate.
This explains why Sol's price hasn't moved. Sol carries the brand, proving that OpenAI remains the company with the highest technological ceiling, while Luna is the revenue engine.
OpenAI hopes developers will adopt a new default habit: use Luna to write agents, run workflows, and perform batch executions—only call upon Sol when facing truly difficult problems.
Once this default is established, competition becomes extremely difficult. Migrating to a new underlying model means retesting, validating, and adapting the entire workflow, and the migration costs will continue to rise.
The true moat has begun to shift from capability leadership to ecosystem stickiness.
Seven: The flagship model no longer determines direction
Looking over a longer time horizon, the AI industry is undergoing a classic inflection point in economies of scale.
When the cost per unit of computing power drops to extremely low levels, total market demand for the token does not decrease with the lower unit price; instead, it explodes exponentially.
The flagship model no longer dictates direction, as the focus of technological advancement has shifted from exploring the limits of intelligence to industrializing compute to reduce costs. When API invocation costs become negligible, the form and boundaries of models will begin to blur.
Enterprise developers no longer focus on the consumption of each token, but instead seamlessly integrate AI into every business process.
The most profound aspect of this shift is that the avalanche of API price cuts is driving up migration costs across the entire software engineering landscape. Once a company’s workflows, agent scheduling networks, and automated data pipelines are all built on combinations of low-cost models, the providers of the underlying models lock in control of the computing pipeline for the next decade.
Over the past few years, large model companies sold intelligence. Starting today, they are selling efficiency.
These are two completely different stories.
[Outside the layout]:
In the past, we were accustomed to comparing AI's development to consumer electronics, anticipating occasional groundbreaking new flagships.
But perhaps the true triumph of AI is not becoming a sensational product.
When the steam engine was first invented, people marveled at its productivity; today, electricity flows everywhere, powering the entire civilization, yet no one ever discusses it specifically anymore.
When humans no longer enthusiastically debate which flagship model has broken another IQ benchmark, AI quietly embeds itself into every system and every command, becoming the default water and electricity behind all automated processes.
Its true era has only just begun.
This article is from the WeChat public account "Beyond the Layout," authored by Hua Hua.
