Thinking Machines, founded by former OpenAI CTO Mira Murati, has launched its first model, Inkling. The model employs a mixture-of-experts architecture with 975 billion parameters, primarily building on the DeepSeek-V3 architecture and leveraging Kimi K2.5 data for post-training. However, in multiple benchmark evaluations, Inkling underperforms compared to Kimi and GLM, yet its usage cost is higher—ranging from $1.87 to $4.68 per million tokens, approximately double that of Kimi. Thinking Machines has raised $2 billion in funding, valuing the company at $12 billion. Murati’s strategy is not to compete directly with China’s strongest models, but rather to target the compliance-focused market in the U.S., where regulatory uncertainty has created demand for a locally developed open model with adequate performance, customization support, and lower compliance risk.Author and source: Letter AI
Former OpenAI chief technology officer Murtaza made a premium alternative to a Chinese model.
Thinking Machines' first model, Inkling, primarily adopts the mixture-of-experts architecture from DeepSeek-V3 and uses data generated by open models such as Kimi K2.5 for post-training cold start. However, in multiple benchmark evaluations, Inkling underperforms compared to Kimi and GLM, while also having a higher invocation cost.
For a star company that raised $2 billion and is valued at $12 billion, with a team that includes many former OpenAI employees, this outcome is somewhat below expectations.
But Murati may never have intended for Inkling to compete for the top spot. As U.S. companies face increasing uncertainty in adopting Chinese models, a domestically developed open model with decent performance, customization support, and lower compliance risk already has its own market. It doesn’t need to be the strongest—it won’t lack for demand.
Is this really Murati’s first model?
On July 15 local time, Thinking Machines Lab released its first model, Inkling, since its founding.
Looking at the specifications alone, this model lives up to the standards expected of a renowned AI lab.
Inkling employs a mixture-of-experts architecture with a total of 975 billion parameters, activating 410 billion parameters for each token processed. It was pre-trained on 45 trillion tokens, encompassing text, images, audio, and video, and supports context lengths of up to 1 million tokens. It can understand text, images, and audio, and outputs results in text form. It also supports adjusting "thinking intensity" to balance performance, speed, and cost.
Inkling has also released its model weights under the permissive Apache 2.0 license. Developers can download the weights and deploy them independently, or fine-tune the model through Tinker, a platform offered by Thinking Machines.
The company has a clear positioning: rather than directly competing with ChatGPT for end users, Inkling is better suited to serve as the foundation for enterprises developing AI applications.
However, the most intriguing part of Inkling is found in the latter half of the technical report.
When introducing its model architecture, Thinking Machines directly stated that Inkling's mixture-of-experts design "primarily follows DeepSeek-V3." It similarly employs a large number of expert modules, activating only a small subset for each token, and adopts DeepSeek-V3's auxiliary-loss-free load balancing design.
The impact of Chinese models on Inkling also extends to the training phase.
Thinking Machines stated that, to initiate post-training, the company conducted a round of supervised fine-tuning using synthetic data generated by open-weight models, explicitly mentioning Moonshot's Kimi K2.5.
The entire technical roadmap is already clear: a U.S. lab founded by the former CTO of OpenAI, whose first model adopted DeepSeek’s architecture and completed its initial training using data generated by Kimi.
There is nothing wrong with learning from open models. DeepSeek and Kimi have released their weights and allowed developers to use them, as scientific research and engineering development have always been built upon the accumulation of publicly available results. More awkwardly, even after Inkling incorporated technologies from Chinese models, its performance still failed to surpass these “teachers.”
Thinking Machines also does not gloss over this. In its announcement, the company directly acknowledged: "Inkling is not currently the strongest model, whether open or closed source."
According to the results announced by Thinking Machines,
In the "Last Exam for Humanity" text-based evaluation, Inkling scored 29.7%, while Kimi K2.6 and GLM 5.2 achieved 35.9% and 40.1%, respectively;
After adding the tool, Inkling rose to 46%, still below Kimi K2.6's 54% and GLM 5.2's 54.7%.
On SWE-Bench Pro, which evaluates real software engineering capabilities, Inkling scored 54.3%, Kimi K2.6 scored 58.6%, and GLM 5.2 reached 62.1%.
The gap on Terminal Bench 2.1 is more pronounced, with Inkling at 63.8%, while Kimi K2.6 and GLM 5.2 reach 71.3% and 82.7%, respectively.
Of course, Inkling also has its strengths. It performs well in web application design, audio understanding, and security evaluation, and can even outperform Kimi and DeepSeek on certain mathematical tasks.
From the perspectives of comprehensive reasoning, programming, and agent capabilities, it is currently difficult to classify it among the top tier of open models.
Performance didn't lead the pack, and the price didn't deliver any surprises either.
On the Tinker platform, the 64K version of Inkling costs $1.87 per million Prefill tokens and $4.68 per million Sample tokens. The former can be roughly understood as the input cost, while the latter is close to the generation cost—and these are already discounted at 50% off for a limited time. The 256K version is priced higher, at $3.74 and $9.36 respectively.
For comparison, the official Kimi K2.6 API charges $0.95 per 1 million input tokens and $4 per 1 million output tokens, while GLM-5.2 charges $1.40 and $4.40, respectively.
Tinker offers both training and sampling services, with billing metrics that are not entirely identical to those of standard model APIs. However, based on the above prices for a rough comparison, Inkling’s prefilling cost is nearly twice that of Kimi K2.6, and its generation cost is also higher than that of Kimi K2.6 and GLM5.2, currently showing no price advantage.
Thus, Thinking Machines' first model presents a somewhat nuanced situation: its architecture draws from DeepSeek, its post-training leverages Kimi, yet its overall performance does not surpass China's leading open models, while its usage cost is higher.
Murati's Choice
Given Murati’s and Thinking Machines’ backgrounds, this choice becomes even more intriguing.
If Inkling had come from an ordinary AI startup that borrowed from DeepSeek and trained with Kimi, it might not be quite so intriguing. But it comes from Thinking Machines—a company that could almost be called an “OpenAI alumni network.”
In February 2025, when Thinking Machines officially launched, the team consisted of only about 30 people, roughly two-thirds of whom came from OpenAI, with the rest primarily from Meta and Mistral.
The founding team once included OpenAI co-founder John Schulman, former OpenAI VP of Research Barret Zoph, and researchers such as Lilian Weng, Andrew Tulloch, and Luke Metz.
These individuals have previously contributed to research in reinforcement learning, post-training, pre-training, safety, and reasoning at OpenAI. At the time, Reuters even directly stated that Murati had recruited at least 20 researchers from her former employer.
Among this group of OpenAI alumni, Murati is undoubtedly the most distinctive.
She joined OpenAI in 2018 and contributed to products including DALL·E, Codex, ChatGPT, and Sora. In 2022, she was promoted to Chief Technology Officer, responsible for coordinating research, model training, safety, and product deployment.
In November 2023, after the OpenAI board abruptly removed Sam Altman, Murti briefly served as interim CEO. Her role was more akin to the overall head of technology and product, giving her a comprehensive understanding of OpenAI’s model development processes and productization pathways.
In other words, Thinking Machines has had OpenAI experience since its inception.
Murali and her team understood how a top-tier closed-source lab developed models and witnessed firsthand how ChatGPT evolved from a research project into a global product. When this group left OpenAI to build their own technology stack from scratch, outsiders naturally expected them to follow a path bearing OpenAI’s imprint.
As a result, Inkling’s most definitive technical origins point to China. This can easily be interpreted as a dramatic signal—that the group most familiar with OpenAI left the company and ultimately chose China’s model as their technical blueprint.
However, this conclusion needs to be examined in more detail.
In recent years, OpenAI’s cutting-edge models have all followed a closed-source approach; external parties cannot access the model weights, nor do they have insight into the full architecture, training data, or post-training methodologies. Even if Murati is familiar with OpenAI’s internal technologies, she cannot directly transfer her former employer’s trade secrets to her new company. In contrast, DeepSeek and Kimi have publicly released their model weights or technical reports, allowing their architectures, training methods, and toolchains to be legally studied and reused.
Inkling is itself an open-weight model. For a company aiming to build an open model ecosystem, adopting mature solutions from open models such as DeepSeek and Kimi is a logical engineering choice.
Today, developing an open model that completely avoids the publicly disclosed achievements of the Chinese team would incur higher trial-and-error costs.
Therefore, Inkling's adoption of the Chinese model does not directly prove that Murati believes DeepSeek's technology is comprehensively superior to OpenAI's. The two parties differ significantly in their levels of transparency and the conditions available for learning from each other.
But this event still sends a clear signal: at least in the open-weight domain, Chinese models have become an important reference point for U.S. AI teams.
Therefore, it's not hard to understand why the Chinese model is being referenced. The real question that needs to be answered is: Why did Thinking Machines, despite absorbing existing architectures and training expertise, produce an Inkling that is both less performant and more expensive?
Murati has a top-tier team, ample funding, and NVIDIA’s latest training systems. She has no reason to overlook the performance and price gaps between Inkling and Chinese models.
In that case, "guì tì" is likely a calculated choice.
"Just average" is enough.
Given the treatment received by Thinking Machines, Inkling should not have been merely a "good enough" model.
In July 2025, before launching any products, the company completed a $2 billion seed round with a valuation of $12 billion, backed by investors including Andreessen Horowitz, NVIDIA, AMD, and Jane Street. Four months later, Thinking Machines was reported to be in negotiations for another funding round targeting a $50 billion valuation; although the deal was never officially announced, the market already regarded it as a potential competitor to OpenAI and Anthropic.
Therefore, when Inkling launched a product that performed worse than Chinese models and was more expensive, it inevitably fell short of expectations.
From a business perspective, Murati may never have intended to use it to compete for the top spot in benchmark evaluations.
Thinking Machines emphasizes open weights, multimodal capabilities, and customizability, hoping enterprises will fine-tune it via Tinker to transform Inkling into specialized models such as customer service agents, programming assistants, and industry-specific agents. Murati is betting on a market that doesn’t prioritize ranking first on benchmark evaluations; businesses care more about whether the model can be deployed privately, whether they can control their data, and whether they can continue training it according to their own business needs.
Previously, a large portion of this demand was absorbed by China's open models.
OpenAI and Anthropic in the U.S. have long maintained closed-source models; after Meta’s Llama 4 underperformed expectations, it has also gradually scaled back its open approach. Meanwhile, DeepSeek, Kimi, Qwen, and GLM continue to open their weights, with performance increasingly approaching that of U.S. closed-source models—at significantly lower costs.
American businesses quickly voted with their feet.
According to data provided by OpenRouter to CNBC, since February 2026, Chinese models have accounted for more than 30% of tokens invoked by U.S. companies through the platform each week, peaking at 46%. In the first half of 2025, the average share was only 4.5%.
During the development of Composer 2, Cursor tested multiple foundational models and ultimately selected Kimi K2.5, as it demonstrated the strongest performance in evaluations. Bridgewater Associates fine-tuned Alibaba's Qwen using Tinker, resulting in a customized model that outperformed some top-tier proprietary models at a lower cost.
China’s open models have become a shortcut for U.S. companies to reduce AI costs. However, this path is now under pressure.
Meanwhile, the regulatory environment in the United States toward Chinese AI models is tightening. Relevant authorities have issued warnings regarding data security, model origins, and supply chain risks, and some companies using Chinese models have come under investigation. Even though no unified restrictions have yet been implemented, policy uncertainty is already sufficient to influence corporate decisions, particularly for companies involved in government business or handling sensitive data.
This is precisely Inkling's market space.
It is an open-weight model developed by a U.S. company, licensed under Apache 2.0, and supports private deployment and customization. U.S. enterprises choose it to avoid political controversies associated with using Chinese models and to more easily pass compliance reviews by government clients and large corporations.
The true value of Inkling comes from this regulatory certainty.
Murati clearly saw the emerging gap.
Thinking Machines incorporates the public achievements of DeepSeek and Kimi, then transforms them into a model offered by a U.S. company. Inkling doesn’t need to outperform Chinese models on all benchmarks—it simply needs to be an acceptable alternative for U.S. companies unwilling to continue using Chinese models.
This also explains why it can afford to be more expensive. Policy has narrowed the competitive field for Inkling, making the gaps in performance and price less critical. For some U.S. businesses, a slightly weaker, slightly more expensive model with lower risk is sufficient.
From this perspective, Inkling better demonstrates that Murati understands business.
However, the cost of this business model may be borne by U.S. companies.
Companies outside the United States can still freely choose GLM, Kimi, and DeepSeek to build products on stronger, more cost-effective foundational models; U.S. companies, however, may be forced to adopt domestically produced alternatives with slightly lower performance and higher costs. The capability gap in foundational models will continue to propagate to fine-tuned products, ultimately resulting in a structural competitive disadvantage.
Washington originally intended to protect U.S. AI companies through restrictions, but may have instead created a market for “expensive alternatives” like Inkling. For Murati, “mediocre” is good enough. For the U.S. AI industry, this “good enough” means higher costs and reduced competitiveness.
