Author: SemiAnalysis
Compiled by Deep潮 TechFlow
DeepChaohao Summary: At SpaceX’s first earnings call, Musk announced plans to add 6–8 GW of computing power by 2027, implying capital expenditures of $300–500 billion that year—comparable to AWS and Google. Though this plan seems outrageous, it is entirely feasible when viewed through three lenses: inference revenue potential, Microsoft’s demand, and SpaceX’s construction speed. Microsoft is aggressively expanding its computing capacity to capture the opportunity of $100 million in annual revenue per MW of API inference, while SpaceX has already proven twice that it is the fastest data center builder in the world.
Musk stunned the world again during SpaceX’s first earnings call, announcing his gigawatt-scale ambitions: he “conservatively” plans to add 6–8 GW of computing power in a single year by 2027, with actual figures potentially exceeding 10 GW. At a cost of $50 billion per GW, this translates to $300–500 billion in capital expenditures for 2027—on par with our expectations for AWS and Google’s spending—yet astonishing for a company with far lower profitability than its competitors.
As we explained in our Meta compute depth analysis, large-scale compute with short delivery timelines is an extremely rare combination, commanding a massive premium—up to $50 billion per GW per year. However, AI labs can afford this price and are thriving.
Our tokenomics model and simulation engine show that, at realistic performance levels (e.g., tokens per second per GPU), both OpenAI and Anthropic could generate over $100 billion in annual revenue per gigawatt by selling API inference services on GB300 clusters—far exceeding the cost of leasing a GB300 cluster for a year at current new cloud provider rates.
For cutting-edge model companies, offering inference token services yields astonishingly high profits.

We assume an annual cost of approximately $12 billion per GW, a conservative rental price of $3 per GPU hour, and estimate token output using our inference simulator based on state-of-the-art model architectures and our agent programming benchmark AgentX (part of InferenceX, built on real production programming traces). We calculate the final valuation by weighting the token output rate according to input, cache read, cache write, and output token costs, along with realistic workload proportions, resulting in an estimated value exceeding $100 billion per GW per year.
Our inference simulator is built from the ground up, grounded in a fundamental understanding of how modern AI accelerators operate. We have developed performance upper bounds and realistic performance models for how cutting-edge models behave during inference, timing each operation and generating accurate execution traces. This is an end-to-end simulation of actual workloads running on real silicon. We have validated the simulator’s accuracy across a variety of accelerators and workloads, and we continuously improve its ability to predict the performance of future accelerators based on design specifications.


Beyond OpenAI and Anthropic, there is in fact a third company in the world capable of achieving such economic efficiency per GW: Microsoft. By having exclusive access to OpenAI’s models, they can generate identical revenue and profit margins per MW without incurring any training costs. Nadella’s negotiations with OpenAI were highly successful: the revised agreement in April 2026 eliminated the previous 20% revenue share. In short, Microsoft has immense incentive to procure as much MW of compute capacity as quickly as possible. Although they currently allocate most of their data center capacity to OpenAI at $14 million per MW per year, they have an opportunity to improve this arrangement. The potential impact is that Microsoft Azure’s revenue growth rate could accelerate from approximately 42% to over 100% next year. This is a once-in-a-lifetime opportunity, and SpaceX is exceptionally well-positioned to serve this demand.

Although Microsoft's contract with SpaceX for 3 GW at $50 billion per GW per year sounds outrageous, we believe it’s possible for two reasons:
- Microsoft is preparing for large-scale data center expansion. As detailed below, they have signed contracts for 10 GW so far this year, with a total contract value exceeding $300 billion (excluding GPU costs). We expect additional contracts to be signed. It should be noted that these contracts provide capacity for the end of 2027 and 2028, and there remains a near-term gap to be filled.
- There is a 90-day cancellation policy, similar to the agreements between SpaceX and Anthropic and Google, with zero balance sheet risk. Given the revenue opportunity, Amy Hood would easily approve this proposal.
For SpaceX, the next natural question is funding. How does Elon pay for such high capital expenditures without the balance sheet of a top-tier hyperscale cloud provider? We expect a combination of the following two approaches:
- NVIDIA's support, in the form of supplier financing, reduces upfront cash costs. This may be why Elon announced exclusive use of NVIDIA on the earnings call! As our accelerator model has repeatedly explained, xAI/SpaceX has actively evaluated alternative options such as TPUs and AMD—so financial considerations likely led them to abandon these options and focus solely on NVIDIA.
- Operated with cash flow financing backed by industry-leading pricing and the fastest delivery cycle: SpaceX will continue to sell large-scale compute at a 3–5 month delivery cycle—an unmatched supply—priced at $30–50 million per MW per year, enabling capital expenditure recovery in less than a year. We explored this in depth in our Meta compute article.
This implies that SpaceX is on track to achieve $300 billion in ARR by the end of 2027, assuming only 50% of the additional computing power added in 2027 is commercialized, with the remainder allocated to the Grok and Cursor teams for training (excluding modeled inference revenue).

Now let’s dive deeper. We start with Microsoft, which has finally awakened in a remarkable way: last year’s pause has been reversed, and so far this year, it has signed binding contracts for 10 GW. We’ll briefly examine the economic model that achieves $100 million in inference revenue per MW per year. Then we’ll turn to SpaceX, analyzing their data center expansion and the feasibility of reaching over 10 GW by the end of 2027.
Microsoft's 10 GW awakening: seize the opportunity of $100 million in annual revenue per MW.
In December 2024, we were the first to highlight a significant pause in Microsoft’s leasing activity within our data center model. Today, the giant has awoken. Our model tracks leasing activity, new cloud provider signings, commencement of self-built construction, and large binding PPAs and ESAs on a quarterly basis. The chart below shows the output: Microsoft has signed over 10 GW across all these categories, equivalent to approximately $300 billion in new binding commitments.

A key driver of this awakening was their urgent need for computing power to capture the $100 million per MW-year revenue opportunity. Microsoft’s $250 billion agreement with OpenAI in October 2025 translates to approximately 7 GW in our Tokenomics model—the world’s best tool for understanding the nuanced mathematics of dollars to watts. This massive infrastructure-as-a-service deal has severely constrained Microsoft’s computing capacity for other use cases, preventing them from leveraging their OpenAI model access to launch the Foundry API business or for applications like Copilot.
However, the profit margins and revenue per MW for these services are the highest to date. We have explained this in detail in our article on AI value capture.
AI Value Capture—The Shift Toward Model Labs

A day in the AI field feels like a year in other industries. Model releases, software breakthroughs, and hardware improvements are compressing cycles that once took years in other sectors into just weeks. Over the past few months, agent AI has crossed a true inflection point, driving a step-function increase in token value while software and hardware advancements have dramatically reduced token generation costs.
To arrive at these margin estimates, we had to carefully synthesize leaked financial data, InferenceX data, microbenchmarks of all latest accelerators in the industry, papers, blogs, and tweets from open-source labs, among other sources. New data points, such as the leaked DeepSeek investor call (where they stated a GPU payback period of 10 months), confirm that our estimates fall within the correct range; however, we first acknowledge that the lack of granularity is indeed unsatisfying. What you truly want to know, beyond a single figure for the company’s overall inference gross margin, is the gross margin for each (model, accelerator) combination across the entire throughput-vs-latency Pareto frontier. For example, what are the gross margins for serving Opus 5 Fast on Trainium3 versus Fable 5 on TPUv7?
We answer this question using a reasoning simulator, a tool available exclusively to SemiAnalysis consulting clients.
Our AI cloud TCO model has addressed cost-related questions, but revenue has historically been unknown. To resolve this, we developed a simulation framework that emulates real model execution on virtual hardware, powered by a fine-grained performance model tuned across various accelerators and operation types. We run each model under all possible service configurations on simulated XPU, blending real-world and idealized service conditions. This enables us to accurately estimate performance for any combination of software, hardware, and workloads, under well-informed assumptions about model architecture.
Thanks to this simulator, our Tokenomics model now includes high-level revenue-per-MW figures for running OpenAI/Anthropic flagship models on all relevant chips. Workload morphology is clearly a major factor—we simulated over $1 million worth of agent trajectories collected from our own usage, while matching real-world interactivity and TTFT levels observed from first-party endpoint access. As a preview, here are our numbers for Fable 5 on GB200 vs. GB300:


This represents a $100 million per MW opportunity for Microsoft. Given the recent surge in demand for Codex and the corresponding acceleration in OpenAI’s ARR, we believe Microsoft can commercialize computing power at a similar rate by providing OAI model services.
Now, to seize this once-in-a-lifetime opportunity, Microsoft and AI labs need data centers—fast and large. SpaceX has already proven twice that they can build faster than anyone else, but by the end of 2026, their computing capacity will be “only” 2 GW. Can they really build over 10 GW in just one year by 2027?
SpaceX: Building data centers at an astonishing pace
In our Meta compute article, we delve into why Elon once again proves himself a business genius: he recognizes that AI labs' profit margins have significantly increased, and accordingly, he has introduced "value-based pricing" for his GPU clusters, rather than the more common "cost-plus" model.
To keep the machines running, Elon needs to build data centers faster than anyone else. We believe he can do it. What gives us this confidence? We’ve written several times about Elon’s speed—including building 300MW of Colossus 1 in 122 days, constructing 200MW of Colossus 2 in six months, and deciding to build an on-site power plant just one kilometer across the border to bypass permitting processes.
Further speed demonstrations are coming. The Southaven power plant has expanded from 27 turbines (approximately 495 MW) in February 2026 to 69 turbines (1.7 GW) in July 2026.


With the emergence of MiniHard, vertical construction beginning in March 2026 could reach 450–500 MW in approximately five months! This doesn’t mean Elon is building better than others—he’s simply using a different approach.
How can all of this be achieved?
Power distribution equipment and high-power transformers have been sold out for two years? Buy power modules directly from China, bypass high-power transformers, and transmit medium-voltage power straight from the generation side to low-voltage transformers—whose supply is far more abundant. Elon’s company employs the world’s best electrical engineers—no one understands how to balance speed and efficiency better. Most global data center operators prioritize efficiency and quality—this is the only way to secure long-term contracts with hyperscale cloud providers lasting 15–20 years. SpaceX, however, focuses on a completely different trade-off: speed above all else. During periods of acute computing scarcity and extremely high AI token profit margins, a 500MW cluster deliverable in three months with a 90-day cancellation window is among the world’s most scarce and valuable assets. All conventional “quality” metrics become irrelevant. Evidence? Google, the most vertically integrated infrastructure company globally, ultimately chose to partner with SpaceX.
Gas turbine orders are booked out until five years from now? GE Vernova is one such case, but there are many other options—our energy model shows that over 30 gas power equipment manufacturers have secured large orders to serve data centers. As long as you’re willing to put in the effort and consider new suppliers, substantial available capacity exists. Trading activity in the secondary turbine market is surging—for example, all turbines originally scheduled for Oracle’s New Mexico site are now on the market. Secondary market prices are high, but Elon can afford them.
Is labor the ultimate constraint? Then maximize parallel work, simplify debugging processes, and pre-assemble as much as possible. Reports indicate that at its peak, Colossus 2 employed approximately 3,000 construction workers per day—about half as many as other gigawatt-scale data centers under construction. Elon has long accomplished feats with fewer personnel than industry standards—just look at Tesla and SpaceX’s histories.
This provides ample time to build multiple such shells before 2027. And this is his first true greenfield project, so the next one is likely to be even better. Of course, another option is to retrofit existing facilities. As we detailed in last year’s in-depth analysis of xAI, both Colossus 1 and 2 were constructed at an astonishing pace through retrofitting.
xAI's Colossus 2—the world’s first gigawatt-scale data center, a unique reinforcement learning approach, and funding progress

There has been extensive coverage of xAI’s Colossus 1. The construction in Memphis is destined to go down in history: the world’s largest AI training cluster, built from scratch in just 122 days. With approximately 200,000 H100/H200 GPUs and around 30,000 GB200 NVL72 units, it remains the largest fully operational, single cohesive cluster (excluding Google).
Building 10+ GW within a year is an entirely different story. SpaceX would need to identify suitable land across the country that is easy to permit and can access natural gas. However, we believe there are sufficient options to support large-scale expansion. This will naturally rely heavily on on-site gas power generation—see our energy deep dive to understand how it works and why it’s essential.
Behind the paywall, we’ll discuss some sites we speculate Elon might choose.
