Recently, some media reported that Microsoft revoked internal access to Claude Code. Claude Code, an AI programming tool developed by Anthropic, became one of the most popular auxiliary development tools within Microsoft after just six months of internal availability. This led to a sharp increase in token consumption and soaring costs, coupled with inconsistent output quality. After careful consideration, Microsoft halted its use and redirected employees toward its own Copilot CLI.
The phenomenon of token consumption being disproportionate to actual output is widespread among other platform companies. Uber exhausted its entire 2026 AI programming tools budget in just four months; some Amazon employees unnecessarily consume tokens; Meta quietly removed its internal employee Tokenmaxxing leaderboard, no longer encouraging token consumption without measurable output. Everyone is embracing AI, but no one has yet found the right approach; companies are all touting AI-native strategies, but (for now) they see no returns—only ever-growing bills. I call this the “token uneconomy.”
Token inefficiency results from the interplay of multiple factors, including poor internal corporate controls, limited returns on token usage, and inherent architectural design issues with agents—such as repeated skill invocations, internal overhead in long-term tasks, and high coordination costs among multi-agent systems. In the future, these issues may gradually ease as internal controls improve and technical inefficiencies are optimized. However, to achieve a positive net token yield, it is essential not only to reduce token costs on the supply side but also to address the challenge on the demand side: enabling token consumption to generate tangible value across a broad range of industrial applications.
Good quality doesn't come cheap.
Over the past two years, mainstream large models have rapidly evolved, and development companies have adopted diverse product strategies based on their market positioning, leading to changes in API invocation prices (per million tokens). While model performance has significantly improved, high quality comes at a cost—the invocation prices for products within the same tier have also quietly increased, becoming a major factor driving up token consumption costs for downstream users.
Leader's tiered strategy
Anthropic was the first closed-source model provider to recognize that programming is the core use case for token monetization. The primary paying users of large models are developers and enterprise tech teams, who are less sensitive to price and place greater emphasis on the model’s coding efficiency and quality. By securing an early advantage in this commercial use case, companies can achieve token premium pricing.
Therefore, Anthropic has focused its R&D efforts on programming. After establishing a strong advantage in programming capabilities, starting with the launch of the Claude 3 series in early 2024, Anthropic became the first in the industry to adopt a three-tier product lineup—flagship, mid-range, and lightweight—enabling tiered pricing within the same model generation and capturing both premium and mass markets. The Opus series, positioned as the industry benchmark for programming, anchors the premium segment at $15/$75 (price per million input/output tokens, same below); the Sonnet series ($3/$15) offers a high-value option for everyday programming and office tasks; and the Haiku series ($1/$5) targets lightweight, fast interactive scenarios with an affordable price point. This precise tiered structure enables Anthropic to maximize profit extraction across each price segment while safeguarding its market share.
This pricing strategy gives Anthropic, as a technological leader, greater competitive flexibility and operational options. For example, upon noticing a rapidly narrowing performance gap with competitors, Anthropic significantly lowered prices with the launch of Opus 4.5, pressuring competitors’ market share. Similarly, with the release of the next-generation model Mythos Preview ($25/$125), Anthropic introduced a new ultra-premium tier within the Opus lineup, raising the price of its flagship product and reversing the previous trend of continuous price reductions in the high-end segment. Subsequently, Fable 5, built on the same underlying architecture, restricted certain features under the guise of security and was priced at $10/$50—still double the Opus series—to target a broader market. Rather than pricing solely by performance, Anthropic now also prices according to the degree of safety constraints, establishing a three-dimensional pricing strategy that segments capability, risk, and price—reclaiming its premium market position.
This positioning strategy was thoroughly validated between 2025 and 2026. Anthropic’s annual recurring revenue (ARR) surged from approximately $1 billion at the end of 2024 to around $45 billion by May 2026. More importantly, this strategy successfully preserved the market premium as a product leadership leader, leveraging performance advantages to escape the race to the bottom on pricing and completing a value loop where high-quality products command premium prices.
Price manipulation by followers
In contrast, during the early stages of commercializing large models, OpenAI and Google chose diversified paths different from Anthropic’s. OpenAI invested heavily in multimodal projects like Sora in 2024, while Google built an ecosystem strategy around Gemini, spanning search, cloud services, Workspace, and other product lines. Although these investments expanded their technological footprint, the分散ed resources resulted in relatively weaker performance in office and programming scenarios. By the time they realized programming was the primary arena for monetizing model capabilities and attempted to catch up, they had already lost their first-mover advantage.
OpenAI has responded with decisive determination. On one hand, it has refocused on coding and agent capabilities, cutting high-cost projects like Sora; on the other, it has followed Anthropic in building its own tiered product lineup, directly competing on every front, while deliberately widening the price gap between its flagship and lightweight models. The flagship model maintains a premium price to uphold its leading reputation, while lightweight models undercut competitors to capture market share. GPT-5.5’s pricing ($5/$30) aligns with Claude Opus 4.7/4.8 ($5/$25), establishing an equivalent premium price anchor. Meanwhile, its mid-tier models—GPT-5.4 Mini ($0.75/$4.50) and Nano ($0.20/$1.25)—are significantly cheaper than the comparable Claude Haiku 4.5 ($1.00/$5.00), using price to gain market advantage.
Google is the core of the Android ecosystem and has already established a complete commercial loop, requiring more complex relationships and more cautious actions. Gemini must serve enterprise customers on Google Cloud, productivity users on Workspace, and consumers using search products. Even while recognizing the importance of programming, it cannot fully concentrate resources solely on programming and productivity—it must still pursue a multimodal and diversified approach.
Google also followed Anthropic in dividing its products into a flagship Pro series and a lightweight Flash series starting with Gemini 1.5, but with a relatively slower product iteration pace and lower pricing. In early 2024, the flagship model Gemini 1.5 Pro charged only $5 per million tokens for short prompts (<128k), one-third the price of GPT-4o and one-fifteenth the price of Opus 3 at the time. By February 2026, the price for Gemini 3.1 Pro’s million-token output had increased to $12, still significantly lower than GPT-5.4’s $15 and Opus 4.6/4.7’s $25. Moreover, Google took a counterintuitive approach by introducing an ultra-lightweight line, Flash-Lite, beneath its Flash series, slashing pricing to match that of open-source models—a classic strategy of trading lower margins for higher volume.
The delayed official release of Gemini 3.5 Pro, which the market has eagerly anticipated, reflects Google’s internal challenges in balancing performance, security, and ecosystem compatibility. The pricing strategy for the new flagship model has also drawn significant market attention.

Figure 1: Pricing Trends for Flagship Models Pricing for the Claude series and GPT-4o/4.1/5.4 is sourced from official pricing pages; pricing for the GPT-5.5 series and Gemini 3.5 Flash is aggregated from OpenAI/Google platforms and third-party sources; GLM series pricing is based on the overseas Z.ai platform, with actual prices subject to exchange rate fluctuations and dual-pricing structures. Chart by Codebuddy
The secondary/lightweight and open-source/semi-open-source model markets are quietly rising in price amid a surge in demand.
Flagship models compete on performance, while secondary/lightweight models compete on price—this is the natural and correct approach in market competition. Amid intense market competition, the typical expectation is that the central price level will continuously decline. However, the reality is the opposite: over the past two years, the price center of the economy-tier token market, composed of secondary/lightweight, open-source, or semi-open-source models, has quietly risen—and it is precisely through this upward shift that the true floor of token prices has been lifted.
On the surface, this is a fiercely competitive red ocean. Low-cost secondary/lightweight models like Sonnet, Mini, and Flash serve as affordable alternatives to mainstream proprietary models, primarily aiming to capture market share. Meanwhile, open or semi-open models such as DeepSeek, Qwen, and GLM have rapidly emerged, adopting a strategy of flagship-level performance at secondary/lightweight pricing, exerting sustained downward pressure on the proprietary secondary/lightweight model market. By the end of 2024, DeepSeek V3 entered the market at approximately $0.27/$1.10—significantly lower than comparable proprietary models. Shortly after, R1 offered enhanced reasoning capabilities at $0.55/$2.19, directly compressing the pricing space for GPT-4.1 Mini and Claude Haiku. GLM-4 Plus delivers near GPT-4-level performance at just $0.69/$0.35, making it highly attractive to price-sensitive developers. Price competition appears to be the norm in this segmented market.
On the other hand, the release of each new generation of secondary/lightweight and open-source/semi-open-source models has been accompanied by an increase in price floors. For example, Haiku 3.5, launched in October 2024, had input/output pricing of $0.80/$4.00; one year later, Haiku 4.5 increased its pricing by 20% to $1.00/$5.00. Around the same time, the GPT Mini series nearly doubled in price, rising from $0.15/$0.60 for 4o Mini to $0.40/$1.60 for 4.1 Mini. The Gemini Flash series saw a similar increase, with pricing jumping from the ultra-low rates of $0.10/$0.40 for 2.0 Flash to $0.30/$2.50 for 2.5 Flash—a more than sixfold increase in the cost per million output tokens. Open-source/semi-open-source models such as the GLM series also saw price hikes: GLM-5’s pricing in overseas markets rose approximately 67% to 100% compared to GLM-4.7. In Zhipu’s own words, this significant price increase reflects the rapid advancement in technical capability and market competitiveness of domestic models.
The fundamental cause of this phenomenon is the explosive growth in consumption of economy-tier tokens. Most everyday coding tasks, document processing, and automation workflows do not require the capabilities of models like Opus or GPT-5.5; instead, they are handled by models such as Sonnet, mini, and Flash, or by open-source and semi-open-source alternatives. With the growing adoption of AI coding assistants, agent workflows, and enterprise AI applications, the usage volume of these secondary/lightweight/open-source/semi-open-source models has surged far beyond that of flagship models. On one hand, this has led to a rapid increase in economy-tier token consumption, making the unsustainable practice of burning cash to maintain low prices untenable; on the other hand, it has created room for vendors to raise prices while demand continues to grow rapidly. As a result, even in the economy-tier token market, the competitive logic has shifted from “which token is cheaper” to “which token offers better value.” Pricing benchmarks have risen across the board—for models such as Claude Sonnet/Haiku, GPT mini/nano, Gemini Flash, as well as DeepSeek, Qwen, and the GLM series.
From the above analysis, it is evident that the token market is undergoing a broad upward shift characterized by high-end pricing power becoming entrenched, mid-tier volumes and prices rising together, and economy-tier players following suit with price increases. Anthropic has established the strongest pricing power in the industry through its leading coding capabilities, while OpenAI and Google are accelerating their pursuit but still need to trade lower prices for volume in the short term. Meanwhile, open-source and semi-open-source models are continuously raising the pricing floor while also beginning to share in the market’s growth红利. This evolving landscape will profoundly impact profit distribution and competitive dynamics across the entire AI industry. As token consumption surges and unit prices rise, the corresponding revenue boom for model providers inevitably leads to higher costs for downstream token users—the fundamental reason why end-user token consumption is becoming economically unsustainable.

Figure 2: Pricing trends for secondary/lightweight and open-source/semi-open-source models. Pricing for the Claude series and GPT-4o/4.1/5.4 is sourced from official pricing pages; pricing for the GPT-5.5 series and Gemini 3.5 Flash is aggregated from OpenAI/Google platforms and third-party sources; GLM series pricing is based on the overseas Z.ai platform, with actual prices affected by exchange rate fluctuations and dual-pricing structures. Chart by Codebuddy
Invisible consumption of agents
While rising token prices certainly strain budgets, what’s even more frustrating is the systemic waste of tokens when calling agents to perform tasks. Context traps, tokenizer black boxes, skill redundancy, and the communication tax and entropy drift inherent in multi-agent collaboration—these structural inefficiencies collectively form the underlying technical causes of token inefficiency.
Context trap
Model inference requires computing relationships between each token and all other tokens, so longer contexts increase computational load and token consumption. The same question, when sent to an Agent without context or history, uses very few tokens. However, if the input includes conversation history, tool logs, code files, error messages, and multi-turn discussions, token consumption may increase by several orders of magnitude.
The agent architecture inherently amplifies the long-text trap. Agents decompose problems, plan tool usage, read files, check feedback, revise plans, invoke tools again, and repeat this cycle—each step potentially reintroducing historical context. The same information is repeatedly read, and the same tasks are repeatedly billed. Salim et al. (2026) found in their analysis of the ChatDev framework that the code review phase consumes an average of 39.5% of total tokens—the highest among all development stages—meaning nearly 40% of token usage is spent on repeatedly passing existing information between agents, rather than generating new content.

Figure 3: Analysis of Token Consumption Proportions Across Stages in 30 Tasks of the ChatDev Framework. Salim, et al., (2026). Tokenomics: Quantifying Where Tokens Are Used in Agentic Software Engineering. Proceedings of the Mining Software Repositories Conference (MSR).
Tokenizer black box
The tokenizer is fundamental to large model training, determining the upper limit of information density for a given parameter count, the lower limit of effective context length, and the reliability in edge cases (such as numbers, code, and multilingual text). The more reasonable the tokenization, the more efficient and stable the model's training and inference become. For open-source or semi-open-source models, the tokenizer and weights are typically public, whereas for closed-source models, the tokenizer is a "black box," and updates to the tokenizer often coincide with changes in token density.
In April 2026, Anthropic released Opus 4.7 and simultaneously replaced its underlying tokenizer. According to Anthropic’s official documentation, the tokenizer adjustment was primarily driven by practical model training needs, adopting a more fine-grained subword segmentation scheme to enhance performance. A side effect is that for texts of the same length, the number of tokens increased by 1.0 to 1.35 times. Independent testing organizations have found even higher actual inflation rates. Finout, an enterprise AI cost management platform, conducted weighted real-world tests on enterprise prompts and found an average inflation rate of 1.47 times (+47%) for technical documentation and English-heavy code files. ClaudeCodeCamp’s comprehensive test across seven real-world file types yielded an average inflation rate of 1.325 times (+32.5%). Developer Simon Willison, through direct API comparisons, found that the same system prompt increased from 5,039 tokens to 7,335 tokens (+46%) under the new tokenizer, while token inflation for high-resolution images reached as high as 3.01 times (+201%).
Previously, when OpenAI released GPT-4o, it upgraded its tokenizer from cl100k_base to o200k_base, nearly doubling the vocabulary size. According to official documentation, this change aimed to improve compression efficiency and enhance multilingual processing capabilities. However, an expanded vocabulary does not necessarily result in fewer tokens for the same text; in fact, for non-English content—particularly CJK characters such as Chinese and Japanese—the altered tokenization granularity of the new tokenizer may lead to an increase, rather than a decrease, in token count.
There is currently a lack of systematic, publicly available justification from model vendors regarding whether finer-grained tokenization improves model performance. In Anthropic’s Opus 4.7 changelog, the new tokenizer is listed under Breaking Changes, describing only the factual change (finer-grained subword segmentation) without elaborating on the technical rationale or performance benefits. Some researchers in the community have noted that finer tokenization could theoretically enhance the model’s vocabulary representation, particularly benefiting code understanding and structured data processing (such as JSON and XML formats, which reached the highest 1.35x expansion limit in Opus 4.7). However, whether this potential performance gain justifies a nearly 50% increase in cost remains an open question.
The tokenizer update frequency is significantly lower than model updates, yet it directly affects the most fundamental billing metric for tokens, and these changes are buried in technical details, making them nearly impossible for average users to detect. Closed-source models are even more opaque about tokenizers, potentially contributing to inefficient token usage.
Meaningless invocation of skills
Skills are one of the key tools that enhance the professionalism of Agent architectures. Some view skills as extended Markdown, others as folders containing various reference materials and operational instructions, while some interpret them as lengthy, structured prompts. In practical reasoning and Agent tasks, many skills are overly long and complex, increasing token consumption.
Gao et al. (2026)’s large-scale empirical study of 55,315 public skills revealed how inefficiently loaded skills waste tokens. At the routing level—where the Agent decides whether to invoke a skill—26.4% of skills had no routing description whatsoever, akin to tool manuals without indexes, significantly increasing the likelihood of ineffective loading. At the content level, over 60% of skill descriptions consisted of background explanations or example text rather than directly executable instructions, meaning most tokens spent on using these skills were consumed reading manuals instead of performing tasks. More seriously, some skills heavily reference external files, injecting tens of thousands or even over a hundred thousand tokens per invocation, with only a small fraction relevant to the current task.
Han et al. (2026) further confirmed the limited utility of skills through the SWE-Skills-Bench benchmark. The study evaluated 49 publicly available software engineering skills on real GitHub projects and found that 39 skills (79.6%) provided no improvement in pass rates—pass rates were identical with or without these skills. The average utility gain across all 49 skills was a mere 1.2 percentage points, while token overhead increased by up to 451%. Only seven skills, which encoded specific domain expertise (e.g., financial risk control formulas, cloud-native traffic management, GitLab CI patterns), delivered meaningful performance improvements (up to 30 percentage points). Additionally, three skills caused performance degradation due to version conflicts (up to 10 percentage points). This demonstrates that skill utility is highly dependent on scenario alignment; indiscriminate invocation only increases costs without benefit.
Multi-agent nonsense and task drift in long tasks
Multi-Agent systems are currently a favored approach, enabling users to lead a team of AI agents—each specializing in coding, reviewing, testing, and fixing—working independently while monitoring one another, often improving output quality. However, agents can also hold ineffective meetings, repeatedly reiterating task context, prior conclusions, and formatted phrases; each repetition consumes additional tokens. Salim et al. (2026) refer to this as the “communication tax” of multi-agent systems.
In addition, delegating complex long-term tasks to multi-agent systems is becoming a mainstream approach in programming and office work, and is gradually expanding into everyday scenarios such as dining and transportation. Long-term tasks are inherently prone to deviation. The context of such tasks is often filled with tool outputs, errors, drafts, and logs, making it easy for model reasoning to gradually drift from the intended goal. To correct this, developers frequently implement additional mechanisms such as summarization, memory, verification, and rollback, which increase token consumption. Luo et al. (2026), in their study of TabTracer, observed that traditional chain-of-thought reasoning tends to enter cyclic states when paths become too long; adversarial injections can deliberately trigger these cycles, causing agents to repeatedly consume tokens on erroneous paths without realizing it. This additional overhead required to maintain stability is commonly referred to as the “entropy tax”—the more complex the system, the greater the autonomy of the agents, the more supervision is needed, the longer the task, and the larger the context, the faster the entropy tax grows. In what appears to be an efficient agent team, more than half of the token costs may be spent on internal coordination and self-correction.
Context traps, tokenizer black boxes, meaningless function calls, verbose writing, and long tasks going off-track—when combined, these factors do not simply add up in terms of token consumption; they multiply exponentially. More notably, these technical inefficiencies affect different users asymmetrically. Developers with technical expertise can mitigate these issues to some extent by adjusting system prompts, trimming skill content, and implementing context window management strategies. However, non-technical enterprise users neither understand the internal token flow mechanisms of agents nor can effectively intervene in their behavior patterns; they only see the numbers on their bills continuously rising, without knowing where the money is going or why so much is being spent. In this sense, token inefficiency is not merely a technical performance issue—it is a question of technological equity. The barrier to using AI tools has shifted from whether one can write code to whether one understands the cost dynamics of agent architectures. In reality, most users of intelligent agents lack such technical background and are placed at a structural disadvantage.
Find genuine demand
Compared to supply-side issues such as pricing and inefficient consumption, the limitations on the application side are a more significant cause of token uneconomics. Although model performance has made remarkable progress over the past two years, the generality of tokens remains quite limited. Current token usage is largely confined to highly digitized scenarios, such as programming assistance, document processing, and data analysis. Beyond these strengths, large model performance declines sharply as the level of digitalization in the application context decreases. In low-digitization offline service sectors—such as food and beverage, housekeeping, retail endpoints, and on-site repairs—tokens can only independently handle tasks within already highly digitized process management areas and struggle to actively participate in on-site operations.
This does not mean AI can never enter these fields, but rather that there is a structural gap between the current pure language model paradigm (token-in, token-out) and the real world. This issue has existed since the mobile internet era and is the fundamental reason why digital technologies have failed to transform the primary and secondary industries. Advances in artificial intelligence are now offering new possibilities for bridging this gap, with foundational research in scientific AI, world models, and robotic systems making progress. Over the past two years, the Nobel Prizes in Physics and Chemistry have been awarded to AI scientists, and humanoid robots from Figure, Tesla Optimus, and Unitree have achieved significant advancements. However, these cutting-edge fields remain in the laboratory stage, and until transformative breakthroughs occur at the application level, tokens will likely remain confined to highly digital environments.
Programming is a special case of generalization.
Programming is currently the application scenario where large language models perform best, but this scenario is not broadly representative; a more accurate description is that it is a general-purpose exception.
Generality means that the programming output is in a universal language for agents, enabling direct orchestration of various types of agents to accomplish diverse tasks in environments with strong digital foundations—where processes and documents are already digitized and algorithmically driven. From this perspective, it is no accident that Anthropic’s Claude Code, focused on programming, and OpenAI’s GPT Codex have become the most popular agent products on the market today.
Special cases refer to programming scenarios where models gain significant advantages during post-training. First, there is deterministic signal feedback: when the model generates code, compilers, interpreters, and unit tests immediately provide precise, structured, and unambiguous judgments of correctness. Second, based on this automated feedback, an efficient closed-loop post-training system can be established, seamlessly integrating feedback into the reinforcement learning loop, enabling the agent to rapidly generate code, encounter errors, and self-correct within a digital sandbox. Such autonomous training environments are rarely seen—and essentially impossible to achieve—in other domains.
Once outside of programming, the efficiency of model training drops significantly. In traditional business domains with relatively low digitization and no ability to form automated post-training feedback loops—such as management decision-making, legal negotiations, clinical medicine, and supply chain logistics—the costs of data collection and result validation consume any token-based economy. Without access to low-cost feedback signals, agents cannot achieve exponential self-improvement and struggle to replicate their massive success in programming.
In February 2023, A&O Shearman (formerly Allen & Overy) became the first law firm to enter into an exclusive strategic partnership with Harvey AI, a vertical large model company in the legal field, deploying Harvey’s AI legal assistant across its 43 global offices. During a several-month trial period, more than 3,500 lawyers at A&O Shearman submitted approximately 40,000 queries to Harvey, covering various legal workflows such as contract drafting, regulatory research, and due diligence, significantly improving operational efficiency.
On the other hand, A&O Shearman explicitly stated in its official press release that all outputs generated by Harvey AI must be carefully reviewed by licensed attorneys before use. AI does not truly replace the professional judgment of lawyers; it merely adds an AI preliminary review step to the existing workflow. When senior partners reviewed contract drafts annotated by AI, the time spent on verification was nearly equivalent to the time required to review the original contracts from scratch. Of course, the feedback from human reviews serves as high-value data for subsequent model training, but the cost of such feedback is significantly higher than that of automated, programmatic闭环. It cannot be ruled out that, in the future, once feedback data accumulates to a certain critical threshold, the agent’s performance in real-world scenarios will improve dramatically, approaching or even surpassing professional levels. However, compared to programming, reaching this critical threshold still has a long way to go.
The difficult leap to the physical world
The primary content of legal tasks remains extensive text processing, a scenario with a high level of digitization that is certainly destined for further digital transformation. As the proportion of tasks that can be digitized and directly controlled and manipulated from the digital world decreases, the percentage of tasks that agents can accomplish will also decline. Although most physical infrastructure is software-driven, relying solely on agents to write code to control the physical world still faces significant obstacles.
Taking the development of humanoid robots as an example, although they have already surpassed the best human marathon times, humanoid robots still struggle significantly with most real-world tasks. Actions that are effortless for humans—such as cleaning, carrying objects, opening doors, and navigating cluttered environments—remain enormous challenges for robots. This is why Moravec (1988) stated, “It is comparatively easy to make computers exhibit adult-level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility.” Nearly four decades later, the insight of this statement has only grown more profound. In her lengthy essay “From Words to Worlds,” Li Fei-Fei identifies spatial intelligence and embodied intelligence as medium-term goals that will require significantly more time to mature. The reason is that the real world has no compiler; the physical world does not accept iteration—it only accepts validation, and the cost of validation is always higher than the cost of generation.
Although simulation technology was once highly anticipated and has achieved some success, there is still a long way to go to match the adaptive performance of agents in programming scenarios. Simulation technology was developed to overcome the challenge of lacking a compiler in the physical world, creating a virtual validation environment using digital twins and physics engines. However, embodied intelligence has encountered the sim-to-real gap: optimal control trajectories trained in simplified sandbox environments using massive amounts of tokens become extremely fragile when exposed to real-world factors such as friction, material fatigue, and environmental noise. Aljalbout et al. (2025) argue that the sim-to-real gap is not a single issue but rather the cumulative effect of multiple sub-gaps—including dynamics discrepancies, perceptual distortions, actuator nonlinearities, and system design flaws—making a perfect simulator computationally infeasible.
In addition, simulation-based training strategies often achieve inflated performance metrics by relying on inaccurate but well-defined boundary conditions in modeling. However, when deployed in real-world environments, these strategies frequently prove unreliable and may even introduce risks. For example, OpenAI’s Dactyl dexterous hand project accumulated the equivalent of 13,000 years of training experience in simulation using 64 NVIDIA V100 GPUs and 920 servers with 32-core CPUs, enabling the robotic hand to achieve extremely high success rates in manipulating blocks. Yet, when confronted with real-world variations in material, temperature, and wear that were not pre-specified, the hand’s robustness rapidly declined. In 2021, OpenAI disbanded its entire robotics team. Co-founder Wojciech Zaremba explained the decision by stating that resources needed to be redirected toward areas where progress is more attainable. Although the company did not officially cite the Sim-to-Real Gap as the primary reason, the industry widely believes that the conflict between the high computational costs of simulation training and the uncertainties of real-world deployment was one of the key factors behind OpenAI’s decision to abandon its robotics initiative.
Validating model performance in the real physical world incurs time and capital costs orders of magnitude higher than in virtual environments, yet such real-world testing is irreplaceable. This asymmetric cost of validation underscores the unique nature of programming scenarios: algorithms are not a panacea, and neither are tokens.
If the practical application of tokens remains long-term confined to programming and a few digital scenarios, and fails to bridge the gap between the digital and physical worlds, the sustainability of AI industrialization and industry-driven AI adoption becomes highly questionable. The future of token economies depends on our ability to extend the effective reach of tokens from digital silos into the broader real world. Before genuine demand emerges in the physical world, the economic inefficiency of tokens may persist for a long time.
The spillover risk of token ineconomics
Token uneconomics is unevenly distributed across the AI industry chain. Upstream infrastructure and hardware manufacturers have reaped substantial profits amid the current surge in fixed asset investment; midstream model providers are still competing on product performance, with high capital expenditures squeezing cash flow; downstream application outcomes vary by user and scenario, with most companies remaining on the sidelines, holding off on investment. Risk in the industry chain is concentrating in the midstream, where model providers are establishing circular financing ecosystems within capital markets. Should the accumulated risks of token uneconomics erupt, they are certain to ripple through financial markets and potentially impact social stability.
Uneven distribution of industrial chain risks
The Token-Agent boom has driven massive capital investment into upstream data centers, networks, and chip manufacturing, as well as power and energy infrastructure. TSMC’s capital expenditure for 2026 is projected to reach $52 to $56 billion. Microsoft, Alphabet, Amazon, and Meta are collectively investing over $300 billion in AI infrastructure between 2025 and 2026, with spending expected to approach $700 billion. Midstream large model providers are the engine of this AI investment wave, the anchor for all AI-related optimism, and the “hope of the entire village.” However, despite explosive revenue growth, major vendors remain deeply unprofitable due to persistently high compute procurement costs. OpenAI anticipates it may not turn profitable until around 2030. Meanwhile, downstream enterprise users who are actively deploying Agents and consuming tokens are already tightening costs. With no clear return on investment yet, setting budget caps for token usage, conducting cost attribution, and restricting access permissions are natural and logical management actions.
We compared the changes in free cash flow (FCF = operating cash flow - capital expenditures) and net profit margins over the past year for representative companies across the upstream and downstream segments of the AI industry (Figure 4). In 2025, upstream companies TSMC (44.5%) and NVIDIA (55.6%) not only achieved higher net profit margins but also recorded robust free cash flow growth of 14.5% and 58.8%, respectively. In contrast, downstream companies Amazon, Microsoft, and Meta, despite maintaining or even improving their net profit margins compared to previous years, saw free cash flow decline by 76.6%, 14.8%, and 3.4%, respectively—primarily due to significant increases in capital expenditures. The gold mine of tokens has yet to be discovered; those digging for gold are still investing heavily, while those selling shovels have already reaped substantial profits.
This scenario has replayed multiple times in history. At the early stage of the industrial revolution, as new technologies emerged, demand first surged among investors and upstream industries. Massive capital expenditures in the midstream generated huge profits upstream, while downstream final consumption was still in its infancy and insufficient to support midstream capacity expansion. Risks concentrated in the midstream, with capital and capacity outpacing actual paid demand. In the short term, valuation corrections, idle capacity, and the exit of some participants are nearly inevitable; in the long term, as long as underlying demand eventually materializes, prematurely built data centers, chips, and networks will still find their purpose, becoming the foundational productive infrastructure supporting economic growth. For the general public and regulators, it is essential to prevent industrial chain risks from propagating outward through financial markets and to mitigate economic volatility caused by risk spillovers.

Figure 4: Comparison of Free Cash Flow Growth and Net Profit Margin Across the AI Industry Value Chain (FY2025–2026) Data Source: Company Annual Reports, 10-K SEC Filings. Chart by Codebuddy
Recurring financing and shadow credit
Industry risks are concentrating among midstream model providers, while some of these midstream firms engage in circular financing with upstream hardware companies, making it unclear whether the growth is genuinely technology-driven or merely sustained by a self-reinforcing capital loop. For example, the “AI perpetual motion machine” formed by OpenAI, NVIDIA, and Oracle: OpenAI first accepts a strategic investment from NVIDIA (originally pledged at $100 billion, later reduced and restructured as participation in OpenAI’s new funding round); OpenAI then uses the raised funds to purchase cloud services from Oracle (with a five-year contract worth approximately $300 billion in compute capacity); finally, Oracle uses OpenAI’s payment commitments as credit enhancement to issue bonds and raise capital to buy GPUs from NVIDIA for data center construction, completing the financial loop. Each step appears to have a plausible business rationale, yet each feels excessively forward-looking.
OpenAI’s total power procurement framework has exceeded $1 trillion, which is inconsistent with its current annualized revenue of $33 billion (as of ARR in May 2026) and is entirely based on expectations of future high growth. If downstream token consumption fails to generate exponential revenue growth for model providers, these “commitments” will turn into a “bubble.” The outlook for token consumption appears pessimistic: according to Bain & Company, to absorb the additional 200 GW of computing power expected by 2030, terminal consumption would need to generate approximately $2 trillion in new annual revenue. Even accounting for cost savings driven by AI, there remains a shortfall of about $800 billion.
Similar cycles of financing games occurred during the internet bubble at the turn of the century, but today, half of the valuation bubble is hidden within the opaque private credit market, making it harder to fully grasp the underlying risks. The Federal Reserve’s interest rate hikes have increased borrowing costs in high-risk markets such as startups and leveraged buyouts, while banks, under Basel Accord requirements, have been forced to exit this market—creating space for private equity firms and ultimately giving rise to a U.S. private credit market worth approximately $3 trillion.
Asset managers such as Apollo, Ares, Blue Owl, KKR, and Blackstone provide 20- to 30-year leveraged financing for data center development through BDCs (Business Development Companies) and direct lending. These loans are typically negotiated privately and priced using mark-to-model methodologies, which can lead to maturity mismatches—aligning cash flows over 30 years with technologies like LLMs that iterate monthly. Additionally, due to model vendors’ cash constraints, interest is often paid in kind (PIK), with payments rolled into the principal. This layers risks that are difficult to detect.
A report from the Bank for International Settlements states that the upside potential of the AI industry chain has already been fully priced into both primary and secondary equity markets, but the downside risks have not yet been priced into the debt market. Should downstream demand emerge slowly and revenues fall short of expectations, the valuation logic underpinning cyclic financing will collapse (equity compression), forcing a reassessment of models in private credit (credit impairments), significantly increasing the risk of a simultaneous collapse in both equity and debt markets.
Resource hunger squeezes out other needs
The expansion of computing power driven by token consumption creates intense demand for resources such as water and electricity, often resulting in significant supply shortages in the short term and putting pressure on local residential water and power supplies.
Northern Virginia’s Data Center Alley hosts the world’s densest concentration of data centers, handling approximately 70% of global internet traffic. Due to long-term wholesale power purchase agreements secured by tech companies, residential and traditional commercial energy allocations have been severely constrained. According to a December 2024 report by Virginia’s Joint Legislative Audit and Review Commission (JLARC), data center electricity consumption has already exceeded twice the output of Virginia’s largest nuclear power plant; meeting the energy demands of only the planned or under-construction data centers in Loudoun County will require adding the equivalent of several nuclear power plants to the grid by 2030.
The intense competition by data centers for high-voltage transmission lines and clean energy has forced local utilities to invest billions of dollars in grid upgrades. Dominion Energy plans to invest billions of dollars over the next fifteen years to expand its grid. These massive infrastructure costs will ultimately be passed on to residents through monthly bills in the form of grid maintenance fees, capacity charges, and other surcharges. Capacity auction prices in Dominion’s service area have surged from $29/MW-day to $444/MW-day—an increase of over 1,400%—directly reflecting severe shortages in power generation and transmission capacity. An analysis by the Piedmont Environmental Council (PEC) of Dominion Energy’s Integrated Resource Plan (IRP) shows that, over the plan’s timeframe, the average residential electricity bill could double.
The crowding-out effect of computing power expansion on everyday demand is not limited to Virginia; major global computing hubs such as Dublin, Ireland; Jurong, Singapore; and Guizhou in China have also experienced similar conflicts. In this sense, token uneconomics is not confined to the digital world—it casts a long shadow over real-life realities as well.
Find the token value equation
Token is one of the most fundamental production factors in the intelligent era. Like other production factors—such as land, data, capital, and labor—misallocation of resources and waste of inputs inevitably give rise to what can be termed “uneconomic” outcomes. In this sense, token uneconomics is not merely a temporary phenomenon during the early stages of the AI industry’s explosion; rather, it coexists with token economics and persists throughout the development of the intelligent economy. At present, token economics has not yet fully emerged, making token uneconomics relatively more pronounced.
Presence does not imply complacency; efforts can be made on both supply and demand sides to reduce token uneconomics and strengthen token economics, ensuring that technological advancements truly translate into tangible economic value. On the supply side, precise technical methods can lower per-token costs, plug inefficiencies, and prevent risk propagation. On the demand side, continuously identifying new use cases can help tokens generate real value. When the supply-side cost reduction curve intersects with the demand-side value growth curve, the net benefit after offsetting token uneconomics against token economics can turn from negative to positive.
Refined technical transformation
Context caching and semantic compression. Context caching has become a standard practice among model providers, significantly reducing input token costs when multi-agent pipelines frequently hit historical caches. However, this approach has limitations: in complex enterprise deployments, cache dispersion caused by highly branched agent paths results in relatively limited actual cost savings. A more fundamental solution lies in context compression—not through simple sliding truncation of historical information, but through active semantic compression that preserves critical instructions and reasoning chains while eliminating redundancy and repetition. This semantic context compression can substantially reduce input token consumption while maintaining instruction adherence.
Skill optimization and subtractive thinking. The SkillReducer study by Gao et al. (2026) identifies two pathways for skill optimization: first, description compression—adding concise information to skills lacking route descriptions while trimming redundant background explanations and examples; second, progressive loading—loading skills incrementally rather than injecting the full skill set into context at once, achieving a 39% reduction in skill payload size. When combined, these approaches significantly reduce token consumption during skill invocation while simultaneously improving model performance by 2.8%. This demonstrates that more skill invocations are not necessarily better; sometimes, subtraction yields far greater benefits than addition. Reducing irrelevant information in the context not only lowers token usage but also enhances the accuracy of model outputs. Here, “less is more” aligns not only with the elegance of code but also with greater token efficiency.
Model routing and task offloading. Using large models for simple tasks is one of the main causes of token waste. By implementing adaptive model routing based on task complexity, simple and high-frequency subtasks are delegated to lightweight open-source models with specialized capabilities, while expensive frontier models are reserved only for critical decision points. This layered approach significantly reduces the average token cost per task without compromising quality in essential环节.
Hard Budget Constraints and Moderator Architecture for Multi-Agent Systems. Without clear division of labor, budget limits, and defined stopping conditions, multi-agent systems are far more likely to devolve into marathon-style roundtable discussions. The solution lies in designing a moderator architecture within multi-agent collaborative networks that incorporates hard budget constraints and asynchronous arbitration mechanisms. The Monte Carlo Tree Search approach proposed by Luo et al. (2026), which introduces tool validation at intermediate steps, saves candidate states, and enables rollback when necessary, can be elevated from a reasoning level to an architectural level: assign token budget limits to each subtask, and have a moderator agent monitor overall consumption to forcibly terminate inefficient loops before the budget is exhausted. This not only prevents financial runaway but often simultaneously enhances the system’s overall efficiency.
Commercial value anchoring
Token governance and cost discipline. Microsoft has restricted Claude Code, and Meta has removed token consumption leaderboards—major companies have shifted from merely encouraging token usage to emphasizing token output and cost discipline. Quotas, approvals, model routing, cost attribution, and team billing will likely become standard practices in enterprise AI governance. This is an inevitable phase as AI enters production systems: even though AI is a powerful tool for driving innovation and accelerating productivity, costs must be clearly accounted for. The number of tokens used, the verifiable outputs generated, and the amount of rework caused must all be measured. Without measurement, there can be no management; without limits, there can be no discipline. The most advanced companies don’t reward the highest AI usage—they reward completing the most work with the fewest tokens.
Rationing will become the norm. Companies will not supply tokens indefinitely; instead, they will manage them like cloud computing resources, setting budget pools and approval processes. This governance does not oppose technological innovation—in fact, rationing will compel architects to design more efficient agent systems that internalize cost constraints.
Identify real-world scenarios for large-scale commercial applications of tokens. This is the fundamental path to achieving positive net token收益. Programming and agent architectures are only small steps toward a token economy; discovering business scenarios that enable massive productivity leaps is a prerequisite for entering the fast lane of token economic development and generating significant economic value. Currently, there are still few cases in real-world business settings where agent architectures are widely applied and generate substantial returns, and most of these are isolated examples. Generalizable solutions applicable across other enterprises and industries are still under development.
Embodied intelligence and digital twins are among the expansion directions, but the asymmetric verification costs brought by the Sim-to-Real gap must be acknowledged. A more pragmatic approach is to identify intermediate domains within traditional industries that offer weakly deterministic feedback—such as image screening in auxiliary diagnosis (referencing established radiological standards), demand forecasting in supply chains (backtested against historical data), and initial contract screening in legal fields (comparable against clause templates). While the verification costs in these scenarios are not as close to zero as in compilers, they are far lower than pure physical-world validation, making them potential bridges for token economies to transition from digital sandboxes into the real world. OpenAI’s recent resumption of robotics research underscores that, despite its challenges, embodied intelligence remains unavoidable.
Return to ROI
Any investment that creates value exceeding its cost, no matter how advanced the technology, will ultimately be unsustainable. Token inefficiency is not a technological failure, but rather a temporary challenge commonly encountered as technology scales toward mass production. Just as early steam engines during the Industrial Revolution were inefficient and consumed enormous amounts of coal, this did not negate the fact that steam power represented the future of productivity. Through continuous improvements in thermal efficiency and expansion of applications, steam power eventually became the most fundamental force driving the first phase of the Industrial Revolution. Today’s tokens and agent architectures are like early steam engines—noisy and fuel-intensive—but they have already demonstrated potential far surpassing human capability in specific scenarios. Their future development will inevitably involve a series of technological innovations, evolving from crude to refined approaches. The most valuable agents in the future will not be those with the most complex reasoning chains, but those that accomplish tasks using the fewest tokens. When the industry transitions from a phase of showcasing quantity to one that values precision, and when every token consumed must justify its value output, tokens will return to ROI as the golden standard—and the agent era will have found its own equation of value.
This article is from the WeChat public account "Tencent Research Institute" (ID: cyberlawrc), author: Li Gang
