OpenAI and Anthropic are collaborating to shift AI model pricing standards from “cost per token” to “cost per task,” aiming to rebuild the value proposition for premium models amid pressure from low-price competition. Companies like Meta and DeepSeek are squeezing the market with extremely low token prices, prompting enterprises to turn toward open-source alternatives. Third-party data shows that Anthropic’s most expensive Fable 5 model accounts for only 11% of Claude software spending; Ramp asserts it has identified a new upper limit for what enterprises are willing to pay for AI, as performance improvements no longer justify higher premiums. This debate over pricing narratives reflects the AI industry’s transition from pure technological competition to a focus on measuring business value.Article author and source: Wall Street Journal
As the AI model price war intensifies, OpenAI and Anthropic are collaborating to redefine the industry’s fundamental metric for modeling pricing—shifting from “cost per token” to “cost per task”—in an effort to rebuild a value framework for their premium models in the eyes of enterprise customers.
Over the past several weeks, two leading U.S. AI companies have spoken out. OpenAI’s Chief Financial Officer, Sarah Friar, wrote in a July blog post that while low-cost model tokens are inexpensive, achieving good results “may require more attempts, more time, or more human review.” Jonathan Pelosi, Head of Financial Services at Anthropic, stated in a recent interview that token cost is merely a “proxy metric” for measuring how much AI compute has been consumed, and that “for every business, a better proxy metric is the actual cost of completing the task.”
The immediate backdrop to this pricing narrative battle is that companies like Meta and SpaceX are pressuring the market with extremely low token prices. According to data from the enterprise spending platform Ramp, corporate adoption of OpenAI and Anthropic has slowed in recent months, with existing customers increasingly turning to cheaper open-source alternatives.
Third-party data provides a more direct signal: Anthropic’s strongest and most expensive widely available model, Fable 5, accounts for only 11% of its Claude software spending. Ramp concludes that a new upper limit for what enterprises are willing to pay for AI has been “found.”
From "Token" to "Task": Resetting the Pricing Anchor
A token is the unit of data processed by an AI model and a measure of computational work. The industry commonly assumes that the lower the cost per token, the more cost-effective the software is for customers overall. OpenAI and Anthropic are attempting to challenge this consensus.
Sarah Friar wrote:
Lower-cost models may have cheaper tokens, but achieving good results might require more attempts, more time, or more manual review; more capable models have more expensive tokens but can complete the same task in a single attempt.Pelosi’s view echoes this: tokens only measure how much AI is used, not the output produced. He believes the cost per token is “only useful as a proxy metric,” and what companies should truly focus on is the actual cost of completing tasks.
Price Shock: From "TokenMaxxing" to Cost Limitation
The shift in metrics occurs amid intensifying competitive pressure, as multiple vendors are pushing for lower token costs, attempting to undercut companies like Anthropic that offer stronger models at higher prices. For example, DeepSeek’s flagship open-source model, released in April, has input and output token costs that are only a fraction of those of Anthropic’s most advanced offerings.
Gil Luria, Head of Technology and Research at D.A. Davidson, drew an analogy: "What the U.S. lab is trying to convey is that using open-source tokens is like using cheap toilet paper—it can get the job done, but you need more of it, and there are other unpleasant side effects."
Companies that once encouraged employees to maximize AI usage through “tokenmaxxing” have now imposed limits after experiencing “bill shock,” Luria said:
For many of these companies, internal costs have become prohibitively high.The most expensive model's performance isn't worth the price: third-party data reveals the upper limit
Although companies like Anthropic have published their own per-task cost estimates, an increasing number of customers are turning to third-party benchmarking firms such as Artificial Analysis and Vals AI for independent evaluations. Rayan Krishnan, co-founder and CEO of Vals, noted:
When laboratories report their own results, they often use internal benchmarks, making true peer comparisons impossible.On a per-task cost basis, Anthropic and OpenAI’s models remain the most expensive on the market, with Anthropic’s Claude Fable 5 leading the pack; however, both also offer more competitively priced options, such as OpenAI’s GPT-5.6 Luna.
What truly dampened the narrative around high premiums was Ramp’s data: Anthropic’s strongest and most expensive widely available model, Fable 5, accounted for only 11% of Claude’s software spending. “Through Fable 5, we identified the new upper limit of what enterprises are willing to pay for AI,” Ramp wrote in the report. “Beyond this point, higher performance is not worth the cost.”
