Microsoft is evaluating Moonshot AI’s Kimi K3 model for its Copilot AI assistant, a move the company estimates could slash inference costs by up to $600 million. It’s worth noting that the $600 million savings figure is based on theoretical scenarios and has not been officially confirmed by Microsoft. The company has also not announced plans to replace any existing AI partners.
The potential savings come from Kimi K3’s dramatically lower per-token pricing. Input tokens cost $3 per million, compared to roughly $5 per million for comparable OpenAI models. Output pricing sits at $15 per million tokens versus approximately $30 for OpenAI equivalents.
What Kimi K3 actually is
Launched on July 16, 2026 by Chinese AI startup Moonshot AI, Kimi K3 packs 2.8 trillion parameters and a 1 million token context window with native multimodality. A million-token context window means the model can process roughly the equivalent of a 3,000-page book in a single prompt.
Kimi K3 scored 1,679 on the LMArena Frontend Code Arena Elo rating, outperforming its Western competitors on coding tasks.
This isn’t Microsoft’s first integration with Moonshot AI. The company previously integrated Kimi K2 variants into its Azure Foundry services in late 2025 and early 2026.
Demand for K3 has been so intense that Moonshot AI temporarily paused new subscriptions due to compute limitations.
Why this matters beyond Big Tech
The semiconductor angle is significant. Microsoft potentially shifting billions in AI spending away from models that require the most cutting-edge chips could ripple through the entire supply chain. Nvidia, AMD, and their various AI chip competitors are all indirectly affected by which models enterprise customers choose to deploy.
What investors should actually watch
The broader dynamic here is the intensifying competition between US and Chinese AI models. Kimi K3’s price advantage reflects different cost structures, different regulatory environments, and potentially different approaches to model efficiency.
Microsoft still has a massive investment in OpenAI, and evaluating a model and deploying it at scale across a flagship product are two very different things.
