Organized & compiled: TechFlow

Guests: Jordan and Max, analysts at SemiAnalysis
Host: Jordan (Internal conversation at SemiAnalysis)
Podcast source: SemiAnalysis
[Emergency Episode] Moonshot's Kimi K3 Has Arrived! China Has a Frontier Model
Broadcast date: July 18, 2026
Key points summary
Moonshot (Moonshot AI) has launched Kimi K3, outperforming Google and Meta on multiple comprehensive benchmarks to become the third-best model in the world, behind Anthropic’s Fable and OpenAI’s Soul 5.6. Two SemiAnalysis analysts, Jordan and Max, broke down the implications of this release in an emergency episode: a 2.8T parameter model is being offered at the same price point as Sonnet ($3/$15 per million tokens). If its scale is comparable to leading closed-source models, Anthropic’s profit margins of $10/$50 could be mindboggling.
A sharper assessment comes from Jordan: the frontier gap is narrowing, which he attributes to U.S. government restrictions on Anthropic/OpenAI releasing their strongest models, artificially creating a window of opportunity for followers. Open source has not truly caught up. Meanwhile, Western open source is completely vacant—no American company’s open-source model can match China’s fifth-best. Model competition is becoming a race of harnessing tools and geopolitical rivalry.
Highlights of insightful perspectives
Regarding the positioning of Kimi K3
- If you look at the overall rankings across all major benchmarks, today there is a very clear top three: Fable, Soul 5.6, and Kimi K3. They consistently rank above everyone else, including DeepSeek, as well as Google, Meta, and xAI.
- Google should feel particularly embarrassed. Just in November and December 2025, everyone still believed the three AI giants were Google, Anthropic, and OpenAI. Even today, talking to some "old-schoolers," they still think so—but clearly, that’s no longer the case.
- It might be the second-best model in the world, because every time I try to do something serious with Fable, I get rejected and sent back to Opus. While I’m not sure if it’s better than Opus, at least I don’t get rejected.
Regarding the profitability of frontier models
- If Kimi is unlikely to be selling K3 at 3/15 at a loss, then Fable, with similar scale, charging 10/50, should completely alleviate concerns that AI labs are unprofitable ventures. Selling tokens at API pricing could even be more profitable than SaaS.
- From K2.7 to K3, the price increased more than threefold, from 0.95/4 to 3/15. But I don’t think there’s much room left for further price increases, as many tasks can be adequately handled by GLM or MiniMax M3.
Regarding the U.S. government and the gap
- I believe this gap has narrowed squarely due to the U.S. government restricting Anthropic, preventing us from accessing these companies' most powerful models. They have artificially caught up.
- We can only access it when frontier intelligence is permitted. This actually presents an opportunity for players ranked fourth, fifth, sixth, and seventh.
Regarding Western open-source vacuum
- I'm shocked that, with the entire market still this inefficient, we don't have a single American company that can at least match China's fifth-best.
- Even if the government does not ban Chinese open-source models, major U.S. corporations are unwilling to feed their proprietary data into Chinese open-source models. Even if you load weights in an air-gapped data center and the CCP cannot see your data, executives still won’t approve it.
About Harness
- Testing Kimi K3 made me seriously consider Open Code, Hermes, and Pi for the first time. Harness is still fully part of the product.
- Simple things can make me choose this model over that one: Can it be installed on a remote SSH server? Are the keyboard shortcuts easy to use? These details within the harness actually determine where I send my tokens—and thus where I allocate my budget.
About "it's still too early"
- I went to ICML last week, and the week before that, I attended the AI Engineer conference. It was ostensibly an AI conference, but over 80% of the people had never heard of SemiAnalysis. You claim to work in AI, but you haven’t even read SemiAnalysis? We’re still too early.
Is Kimi K3 the third-best model in the world?
Jordan: Quick hot take—is the Kimi K3 now the third-best model in the world?
Max: The answer is a clear "yes." Everyone likes to complain about benchmarks, and benchmarks do have issues, but if you look at the overall rankings across all major benchmarks, their direction has consistently been correct. Today, there is a very clear top three: Fable, Soul 5.6, and Kimi K3—consistently outperforming everyone else, including open-source players like DeepSeek, and significantly outperforming Google, Meta, and xAI. This is an extraordinary achievement for the Moonshot AI team.
Google should feel particularly embarrassed. Just in November and December 2025, everyone still believed the three AI giants were Google, Anthropic, and OpenAI. Even today, talking to some "old-school" folks, they still think so—but clearly, that’s no longer the case.
Overall, I still feel it’s not quite as good as Fable and Soul 5.6. It’s a bit funny that they explicitly stated this themselves in their model release blog. Maybe it’s old-school Chinese modesty, or maybe they didn’t want to attract scrutiny from the U.S. government, especially since the release of Fable 5.6 was somewhat delayed. But regardless, it’s very impressive.
Jordan: They wrote in the blog’s limitations section: "Although K3 is overall a highly competitive model, there remains a noticeable gap in user experience compared to Fable 5 and GPT 5.6." My personal experience is that it’s indeed good, but it’s really slow—and that’s frustrating. It’s the first time I’ve been motivated to try an open-source harness. Honestly, I feel I’ve learned more about harnesses than about the models themselves, because all these models are good enough to handle the basic tasks I’m currently doing; I struggle to find complex tasks they can’t accomplish.
This might be the second-best model for me, because every time I try to use Fable for serious tasks, I get rejected and bounced back to Opus. I’m not sure if it’s better than Opus, but at least not getting rejected is less annoying. However, I never get rejected when using an API key on a pay-per-use basis—only when using the web console or deep research. Clearly, they don’t have enough GPU capacity to meet the demand for this model. In the past, they solved this issue by going open-source and releasing the weights for others to serve. But this time, they haven’t released the weights yet and say they’ll do so in 10 days.
Why delay open-sourcing the weights by 10 days?
Jordan: What do you think the strategy behind this delay is?
Max: To be clear, this is all my pure speculation. One major reason might be that they need to give the vLLM and SGLang inference engine teams enough time to ensure they can serve this model with high performance. If they released the weights today and everyone was serving it but only delivering 20 tokens/second, it would be terrible for their brand capture. They currently have a fantastic opportunity to gain massive PR and adoption—if the launch is hampered by performance issues, it could undermine the momentum.
Another possibility is that they are negotiating licensing partnerships with inference service providers such as Together AI, Fireworks, Nebius, and Groq to use the latest chips like the GB300 to serve incremental capacity. These two factors are likely the main reasons for the 10-day delay.
Profitability of cutting-edge models
Jordan: Let’s talk about the model architecture. With 2.8 trillion parameters, it cannot fit on a B200—you need a B300, GB300, or AMD MI355X to serve it on a single 8-GPU HGX server. You could theoretically use cross-node pipeline parallelism, but that would severely impact performance. Therefore, only those with the latest chips can serve this model.
Going back to what you said about Google, this model already achieves frontier-level competitiveness with 2.8T parameters, which gives us some insight into how large closed-source frontier models might be. If they’re using 10T parameter models to compete with this, it would be even more embarrassing. We have to assume it’s on the same scale as Soul and Fable.
Max: Yes, you're right. I still have full confidence in Anthropic’s research team’s capabilities and insight. If someone on Twitter claims current closed-source models have 10 trillion parameters, and that were true, then buddy, you’d better pack your bags—NVIDIA’s stock would drop 50% tomorrow, and it would all be over.
I’m fairly confident that Kimi K3 isn’t much smaller than today’s leading proprietary models, and may even be slightly larger. If true, this further supports a point we’ve consistently emphasized at SemiAnalysis: the profit margins of these proprietary labs are absolutely mindboggling. If Kimi is almost certainly not losing money by serving K3 at $3/$15—same pricing as Sonnet—while Fable, of comparable scale, charges $10/$50, then any lingering concerns about AI labs not being profitable should be彻底 dispelled. Selling tokens at API pricing may already be more profitable than SaaS, today.
Jordan: No employee costs, just GPUs. How does that compare to the previous pricing? You mentioned 3/15, but the previous Moonshot direct pricing was 0.95/4, so the price increased more than threefold from K2.7 to K3. How much more room do they have to raise prices?
Max: I don’t think they have much room to push prices higher. Even at 3/15, many will find it too expensive. Their use cases are already covered by GLM or MiniMax M3. There’s an interesting split here: companies like SemiAnalysis, who don’t care about burning Dylan’s tokens, will keep using Fable for nearly everything; while extremely cost-sensitive users, like Tesla or Uber, spending only $200 per week on tokens, will opt for the GLM pricing tier. So who exactly will switch to Kimi K3? Probably just a group philosophically passionate about open source and eager to support new models. I wouldn’t be surprised if large enterprises ultimately choose not to adopt this model.
New Architecture and Next Steps
Jordan: This is a completely new architecture. With 2.8T parameters, it incorporates Kimi’s delta attention, potential residuals, and stable latent features—essentially a scaled-up version of the previous model, roughly twice as large. Previously, with K2.5, we saw Cursor use it as a composer, built on continued pre-training and MRL, leading to checkpoints like 2.5, 2.6, and 2.6.7. This is a new base model, yet it’s already highly polished, without the rough edges commonly seen in the original models. What’s next? When will 3.1 be released? Will pricing change? Will there be a composer based on K3?
Max: There probably won’t be a Composer based on K3; the Cursor team has already decided to train their own model from scratch. As for K3.1, K3.2, etc., we can expect two or three updates over the next one to two months, continuing with post-training. I suspect pricing will remain unchanged, since they won’t be running on new hardware in the next few months and thus won’t have improved throughput to justify a price cut. Maybe an exceptionally skilled kernel engineer could bring costs down to DeepSeek V4 levels, but I’m skeptical that a 3T-parameter model could achieve that. Current pricing from GLM and MiniMax may already be at the limit of what’s feasible for 1T to 1.5T-parameter models.
Will open source catch up to closed source?
Max: The more interesting question is whether the gap between open source and closed source will continue to narrow, and whether open source can truly achieve parity with cutting-edge levels. If this happens, it would have a massive impact on our entire industry. What do you think?
Jordan: My view is that the gap has now narrowed, squarely due to the U.S. government restricting Anthropic, which has prevented us from accessing these companies’ most advanced models. They have been artificially caught up.
We can see the comparison between Mythos and Fable. I can’t use Mythos at all, and I have to beg hard to occasionally get access to Fable. As for 5.6 Soul, our internal assessment is that it’s not the largest model trained by OpenAI—it’s smaller than 4.5. They have an even larger one in their possession. As a result, we can only access it when frontier intelligence is permitted by the government.
This is actually an opportunity for the players ranked fourth through seventh to let loose within certain limits and start capturing user share, but they will never reach the true frontier. I think the frontier could make a significant leap by the end of this summer, or it could shift slightly with changes in political direction. It’s also possible we’ll begin to discover modalities beyond coding, enabling genuine exploration of those domains.
By the way, the Inkling release from Thinking Machines is interesting—I find the native audio input to be a signal of the future.
Max: The West is truly, truly, truly lacking a decent open-source model. It’s astonishing that, with this market still so inefficient, not a single American company can even match China’s fifth-best. On one hand, a complete U.S. government ban on Chinese open-source models may just be a matter of time. On the other hand, even if such a ban doesn’t happen, ordinary American enterprises won’t risk feeding their proprietary data into Chinese open-source models. Even if you load weights in an air-gapped data center—where the CCP can’t see your data—executives still won’t approve it. Many companies care about token budgets and only want to run Western or non-Chinese models. Inkling is currently the best Western open-source option we have access to, but it’s still far from the cutting edge—and that shocks me.
Jordan: Previously Neotron, now Inkling. I see two opportunities in Inkling’s strategy: first, they must be better than most Chinese open-source models to even enter the game. Second, they must outperform secondary and tertiary models from leading labs, even surpassing Sonnet, because you can access near-cutting-edge intelligence via Bedrock or Foundry using closed-source secondary models to save costs. I’ve never fully understood the Western open-source angle of “helping people save money.” While advancing models within ecosystems like Fireworks, Together, and Base10 is certainly positive, the bulk of the market lies at the governmental level.
Built for Chinese-made accelerators
Jordan: Another point worth mentioning is that the K3 blog stated quantization was performed during the SFT phase, using MXFP4 and MXFP8 for weights and activations, with the official claim being "broad hardware compatibility." What other hardware do you think Moonshot might care about?
Max: I have a list of 11 Chinese accelerators—you should subscribe to SemiAnalysis’s accelerator model to learn more. Huawei Ascend, Baidu Kunlun, Cambricon, Moore Threads, and various other chips appear in papers and are visible in code. Running cutting-edge models on domestic accelerators has become a national priority in China. If we still call Google a frontier lab by the end of 2025, then we should also call Moonshot a frontier lab now.
Jordan: By the way, my dad is on a business trip in China, and he said all the hotels are fully booked because Xi Jinping is about to deliver a speech in that region stating that AI is China’s top priority. A lot of what you said is correct.
Harness is the product itself.
Jordan: The biggest insight I’ve gained while using these models is that it’s becoming increasingly difficult to distinguish between using the most cutting-edge model with max thinking mode and using a medium-effort approach. In everyday tasks, I genuinely can’t find anything these models can’t handle. My default behavior is to always use the most powerful, most intensive mode, because I don’t care about Dylan’s budget.
But there’s one layer where Harness itself is part of the product. Testing Kimi K3 made me seriously evaluate Open Code, Hermes, and Pi. Harness is still fully part of the product. Simple details can make me choose one model over another: Can it be installed on a remote SSH server? Are the shortcuts intuitive? Can I edit previous commands? These small details within Harness actually determine where I send my tokens—and thus where I allocate my budget.
Max: Many people talk about token budgets, but based on your workflow description, even for tasks that GLM could handle, I’d still prefer routing them to Fable using Max Intelligence because the ROI is worth it. Benchmarks suggest many tasks can be migrated to GLM, yet you’re still happy staying with Anthropic or OpenAI models.
Jordan: Basically, yes. But I use a lot of Slack bots and don’t know what models are running behind them. For example, with Perplexity’s Slack integration, if it starts routing to K3, or to GLM, or to Sonnet, I really don’t care. It was only when I looked at usage that I realized how much was using OpenAI models, because that decision was made by the system itself. That justification is handled by the harness.
Max: This is actually an entry point for outcome-based pricing. If a lab adopts outcome-based pricing, it could achieve gross margins of over 95%, because the tasks you're currently paying for with stable pricing can today be handled for pennies.
Jordan: Second, I don’t believe these labs have run out of ideas. They can continue training astonishing models to beat coding-side RSI, but withhold them from us, maintaining their “permanent underclass.” They can keep distilling, giving us just a taste, while continuing to explore other applications—like video generation, audio-to-audio, deep research—things quite different from coding. Robotics and world models are a straightforward direction; what if Anthropic shifted its focus from knowledge work to physical labor? I don’t believe they can’t build a sustainable, high-ROI business using the world’s greatest technology.
We're still too early
Max: Even without considering these points, I use these models every day far more than my software engineer friends—they use them ten times less and spend ten times less. Someone using Fable uses it just as much as someone using Sonnet, but the Fable user spends ten times more money, generating 90% gross margins and supporting the bulk of the business. Once these users start adopting larger models and using them even more, demand will only grow further—even without the models becoming better. Then I still have to talk to neighbors who aren’t into technology; among them, I’m probably in the 0.1% or even 0.01%, with potential for 1,000x growth. Back to Masa-san’s (Masayoshi Son) “Golden Goose Index Curve.”
Jordan: She’s absolutely right that it’s still too early—that’s why I don’t think Kimi K3 will slow the net new ARR growth for Anthropic and OpenAI. Even if some current users of Fable and 5.6 switch irreversibly to Kimi K3, they’ll be completely overshadowed by those who haven’t yet seriously tried this technology. These users are discovering new high-ROI use cases every day, and they’ll still default to 5.6 Soul or Fable 5 to unlock these new scenarios. You won’t see the ARR growth rate slow down.
Max: Think about how many people haven’t subscribed to this podcast or followed SemiAnalysis. I went to ICML last week, and the week before that, I attended the AI Engineer conference—a so-called AI conference—where over 80% of attendees had never heard of SemiAnalysis. You claim to work in AI, but you’ve never read SemiAnalysis? We’re still too early.
Jordan: This is an ego check for you, Max—take a breath.
