Author: Jamin Ball
Compiled by Deep潮 TechFlow
Shenchao Summary: xAI’s Colossus 1 data center has just found a buyer—Anthropic has secured all the computing power for an annual payment of $5 billion. This deal instantly elevates xAI to the status of a new cloud computing giant, while Anthropic can leverage this computing capacity to generate $15 billion in revenue; both parties stand to profit immensely.
Let’s do a rough calculation! (All estimates on scratch paper...)
Assume Colossus 1 has 220,000 GPUs.
Assume 150,000 H100 units, 50,000 H200 units, and 20,000 GB200 units.
Rental assumption:
- H100 at $2.30 per hour
- H200 at $2.60 per hour
- GB200 at $5 per hour
The combined rental cost for the entire cluster is approximately $2.60 per hour.
Assume all are take-or-pay agreements (you pay for full capacity at 24x365).
This is equivalent to xAI generating $5 billion in annual revenue. We are witnessing the birth of a new cloud computing giant!
And more than that—in a recent podcast with Dwarkesh, Dario roughly calculated the unit economics (he specifically emphasized that this was industry-standard math, not specific to Anthropic—which is crucial, as he did not disclose any information about Anthropic).
He mentioned using $100 billion in compute spending as an example (he simply picked a round number). This spending would be allocated between training and inference. If training takes too large a share, your revenue won’t be sufficient; if inference takes too large a share, you’ll undermine future R&D progress. He believes the industry currently splits compute spending between training and inference at a 50/50 ratio. He stated that, as an industry, that $50 billion in inference spending could be converted into $150 billion in revenue (he specifically noted this is likely the industry’s unit economics model one to two years from now).
Apply this logic to xAI’s transactions. Under the above assumptions, Anthropic pays $5 billion annually. Assume this translates to $15 billion in annual revenue (60–70% gross margin).
Win-win!!
Dario said exactly:
Think of it this way. Again, let me emphasize that these are simplified facts—the numbers aren’t precise. I’m just creating a toy model. Suppose half of your computing power is used for training and half for inference. Inference has a certain gross margin.

@GS_CapSF "We have signed an agreement with SpaceX to utilize the full computing power of their Colossus 1 data center, enabling us to gain over 300 megawatts of new capacity in one month (equivalent to more than 220,000 NVIDIA GPUs)."
