Nine flagship AI models launched in one month, signaling the 'disposable' era

iconMetaEra
Share
AI summary iconSummary
In July 2026, nine flagship AI models entered the market, including Tencent Hunyuan Hy3, Kimi K3, and Qwen3.8-Max. The rapid releases reflect intensified competition and accelerated innovation. Kimi K3 ranked first on Code Arena’s frontend leaderboard, while Hy3 achieved results with fewer active parameters. Pricing and code performance are now as critical as raw computational power. Traders are monitoring how these developments impact the Fear & Greed Index and top altcoins.
The shelf life of the "strongest model" is shortening.

Article author and source: Dingjiao

Within one month, nine flagship models were released in quick succession—an uncommon occurrence in the large model industry.

From the official release and open-sourcing of Tencent HunYuan Hy3 at the beginning of the month to Anthropic’s release of Claude Opus 5 at month’s end, models such as Kimi K3, Qwen3.8-Max preview, and Grok 4.5陆续 debuted. July became one of the rare release peaks in the past year.

Different companies chose different timing. Some released their products following competitors' updates, others unveiled theirs around the World Artificial Intelligence Conference, and some responded to market competition with price adjustments.

But July was not only notable for the high number of model releases; more importantly, it reflected two significant changes.

First, the performance gap between models is narrowing. According to the latest Composite Intelligence Index from Artificial Analysis, the differences among top models have significantly narrowed. With rapid version iterations, it is increasingly difficult for any single model to maintain a long-term lead as companies begin focusing on excelling in specific tasks rather than pursuing absolute dominance in a single dimension.

Second, open-source models are now entering the top tier. Among the new models released in July, Tencent HunYuan Hy3 and Kimi K3 have chosen to open their weights (making model parameters available for developers to download and deploy independently). Notably, Kimi K3 ranks first on the Code Arena front-end programming leaderboard, surpassing GPT-5.6 Sol and Claude Fable 5. Over the past two years, open-source models were largely seen as followers of proprietary models, but this round of competition demonstrates that open-source models are now directly competing with proprietary ones in certain development scenarios.

So, what changes and implications have resulted from this month's intensive releases?

01. No longer just comparing who is smarter

Over the past few years, the large model industry primarily competed on general capabilities—whose knowledge base was broader and whose responses were more accurate. But now, the competitive landscape has shifted.

This round, coding ability has become the most obvious area of competition.

OpenAI continues to strengthen programming and complex reasoning, and is further expanding the Codex ecosystem to bind developers through development tools. Anthropic is betting on code understanding and security auditing to help developers tackle complex engineering challenges in large-scale projects. Grok 4.5 emphasizes speed and low cost, and is integrated with the AI programming tool Cursor to cover everyday development scenarios.

Researcher Yao Yuhang told 'Dingjiao One' that all companies are heavily investing in programming capabilities, and the gaps are rapidly narrowing—for example, Kimi K3’s coding ability has significantly improved and is approaching the level of leading proprietary models.

Image source / pexels

Agents are another frontier being fiercely contested by various companies. GPT-5.6 Sol, in conjunction with Codex, explores automated programming; Opus 5 emphasizes long-term task execution; Kimi K3 demonstrates the ability to run continuously to complete complex tasks; and Tencent Hunyuan Hy3 seeks to integrate with the WeChat ecosystem to explore agent applications.

However, agents are still immature in practical applications. Independent developer Beicheng noted that the current limitations of some agent products are not about whether they can invoke tools, but rather their task scheduling capabilities. For example, Codex’s Ultra mode launches multiple sub-agents that redundantly read the same files and perform similar analyses, leading to increased token consumption without a corresponding increase in meaningful work. In his view, improving agent capabilities cannot rely solely on increasing the number of sub-agents; the key lies in the ability to determine which tasks are worth analyzing and which steps can be omitted.

Yao Yuhang holds a similar view. He believes agents are currently the most promising area to focus on, but they are still far from being fully mature. In his opinion, there remains significant potential for agents in vertical sectors such as healthcare and finance. The key differentiator in the next phase will not only be the model’s capabilities, but also how practical and user-friendly the agents are in real-world scenarios—and whether customers are willing to pay for them.

Domestic models are seeking breakthroughs by enhancing various capabilities. Kimi K3 focuses on strengthening long-context understanding, front-end programming, and product design, ranking among the top in multiple programming benchmarks and approaching the capabilities of overseas proprietary flagship models.

The trade-off between model scale and inference efficiency has also become another key thread.

In July, several models were released with parameter counts reaching trillions and even tens of trillions; however, parameter scale can no longer be simply equated with performance. With the advancement of MoE architectures, models can now possess much larger parameter capacities while activating only a small portion of them during inference, thereby reducing actual computational costs.

The industry has thus developed two paths: one is to continue scaling up, trading larger parameters for greater capability; the other is to emphasize efficiency, achieving performance close to flagship models with smaller activated parameters.

Price also became the most direct variable in the July model competition.

Overseas vendors are more likely to adopt differentiated pricing strategies. OpenAI has further segmented the market by splitting GPT-5.6 into multiple versions, allowing users to select the model best suited to their task complexity. Anthropic continues to focus on its high-performance flagship models, covering high-value development and enterprise use cases through its Opus series. Grok 4.5, on the other hand, adopts a low-price strategy, offering output at just one-quarter the cost of the Opus series, while attracting developers through integration with Cursor.

Domestic manufacturers, on the other hand, emphasize cost advantages. Over the past six months, multiple domestic model providers have continuously lowered their API prices to reduce usage barriers and expand their base of developers and enterprise users.

In July, the competition among large models expanded from single-capability comparisons to multiple dimensions, including code, agents, parameter efficiency, and pricing.

02. Who exceeded expectations, and who received polarized reviews?

Based on comparisons with previous generations, ranking performance, and developer feedback, the models released in July can be broadly categorized into three tiers.

First, let’s look at the ones that exceeded expectations: Kimi K3, Grok 4.5, and Hunyuan Hy3.

The most closely watched model is Kimi K3. Built on a MoE architecture, it boasts a total parameter scale in the trillions, with significant enhancements in long-context understanding, front-end programming, and product design capabilities. On its release day, Kimi K3 ranked first on the Code Arena front-end programming leaderboard with a score of 1,679, rising 17 positions from Kimi K2.6’s previous rank of 18th. One week later, Kimi K3 entered the global Top 10 on the Arena overall leaderboard for the first time.

However, Kimi K3 also has clear capability limits. In the Artificial Analysis knowledge reliability test, its hallucination rate increased from 39% in Kimi K2.6 to 51%, and its output speed is approximately 33 tokens per second, significantly lower than that of similar models.

Developers’ real-world experiences also confirm this “one-sided strength.” Bei Cheng believes that Kimi K3 has made noticeable improvements over its predecessor, with advantages primarily in front-end development, UI design, and product expression. For instance, when optimizing website UIs or designing app interfaces, Kimi K3 delivers superior results, but its performance becomes less consistent when handling complex engineering tasks.

Image source / pexels

Grok 4.5 has also attracted significant attention. After its release on July 9, it ranked among the top entries in the Artificial Analysis Intelligence Index and demonstrated standout performance in certain tool-calling tests, making it a popular cost-effective choice among many developers.

Beicheng’s experience aligns with this—he believes Grok 4.5’s greatest advantage is its speed; for the same task, it can often be completed in three to five minutes, whereas GPT-5.6 sometimes takes twenty minutes or even half an hour. Combined with its lower price, it is well-suited for scenarios requiring rapid development. Yao Yuhang believes that Grok 4.5’s programming capabilities have been specifically optimized, and when used with Cursor, the overall experience is excellent. A relevant background is that SpaceX previously announced the acquisition of Anysphere, Cursor’s parent company, with the transaction expected to close in the third quarter; the data collaboration between xAI and Cursor is also seen as helping to enhance Grok 4.5’s programming experience.

Hunyuan Hy3 represents a breakthrough in the efficiency pathway. With a total of 295 billion parameters and only 21 billion activated parameters, it achieves performance close to much larger competitor models on multiple public benchmarks, despite having less than one-fifth the parameter scale of flagship models. However, on SWE-Bench Pro, Hunyuan Hy3’s score of 57.9 lags behind Claude Opus’s 69.2, indicating that its capability in complex code repair still needs improvement.

Yao Yuhang believes that Hunyuan Hy3 represents an improvement over Tencent’s previous models, and with its open-source nature and compact design, it’s a solid choice among mid-sized models.

Now consider those that remain in the top tier but offer limited surprises: GPT-5.6 Sol and Claude Opus 5.

GPT-5.6 Sol and Claude Opus 5 remain in the top tier, but without significant breakthroughs. GPT-5.6 Sol has enhanced its complex reasoning and agent programming capabilities, achieving 88.8% on Terminal-Bench 2.1 (a benchmark testing a model’s ability to complete real-world tasks in a command-line environment), with Ultra mode further improving to 91.9%. Claude Opus 5 continues to excel in complex tasks, placing greater emphasis on reliability in scenarios such as advanced engineering, code understanding, and security analysis. Opus 5’s advantage lies in performance approaching that of Fable 5, at only half the price.

Both models remain in the top tier, but have not delivered any major breakthroughs.

Bei Cheng mentioned that GPT-5.6 already covers most of his daily development needs, but when encountering particularly complex issues, he turns to Opus 5 for solutions. However, greater capability also means higher costs; for him, Opus 5 serves more as an "expert" to be consulted when addressing critical problems.

Finally, a category with divergent feedback and still pending validation: Gemini 3.6 Flash and Qwen3.8-Max preview.

The most notable example is Gemini 3.6 Flash, which scores 50 on the Artificial Analysis Intelligence Index, matching its predecessor, 3.5 Flash. The developer community generally agrees that this version shows no significant improvement over the prior generation and underperforms compared to competing models released at the same time. However, some consider it acceptable for a lightweight model, as its output quality is generally adequate.

According to independent developer Kapdim, Gemini 3.6 Flash offers a pleasant everyday experience; while its programming capabilities are not outstanding, its multimodal understanding remains quite strong. It supports input in any file format and provides a more natural conversational style and better user experience than many competitors.

After the Qwen3.8-Max preview launched in July, the model's performance proved inconsistent, with mixed market feedback. Bei Cheng noted that the Qwen3.8-Max preview demonstrated deep reasoning but occasionally became stuck in prolonged thought processes, as if "getting lost in its own logic," without ultimately delivering effective solutions. Kapdim felt the Qwen3.8-Max preview fell short of expectations; when testing it using his own frontend aesthetic evaluation set, he found a clear gap between its actual capabilities and the advertised claims. As a preview product still under refinement, its final performance will need to be validated upon the official release.

In July, no single model led across all dimensions; different models began to specialize around specific use cases.

Taking Beicheng as an example, he selects tools based on the task: GPT-5.6 handles daily development, Grok 4.5 is used to quickly fulfill requirements, Kimi K3 is employed for product and frontend scenarios, and Claude Opus 5 tackles complex problems. Kappidim’s approach differs—he uses Claude’s Fable 5 or Opus 5 models for planning, followed by execution via GPT-5.6 Sol.

03. Releases are becoming faster, and the debate between open-source and closed-source is intensifying.

Behind this wave of密集 releases, several questions deserve analysis: why is model iteration accelerating, why is it concentrated in July, and why is open source disrupting existing business models.

First, the approach to model iteration is changing. In the past, improvements in large model capabilities primarily relied on retraining larger models, which was time-consuming and costly. However, as technologies such as reinforcement learning and post-training (optimizing model performance after base model training through fine-tuning, reinforcement learning, etc.) have matured, model upgrades are increasingly driven by post-training optimization.

Yao Yuhang told 'Dingjiao One' that over the past two years, the pace of model releases has significantly accelerated, largely because the industry has mastered more mature post-training methods. Now, many model improvements don’t require training a new model from scratch; instead, they involve further optimizing existing base models through techniques like reinforcement learning and self-iteration.

This means the cycle of model competition is shortening. In the past, a leading model might maintain an advantage for months; now, the competitive edge from a new version may last only a few weeks. If the release pace falls behind, a model can quickly be overtaken by a newer version. For manufacturers, releasing a model is no longer just a technical demonstration—the timing of the release itself has become part of the competition.

Image source / pexels

Secondly, July is a period when multiple milestones coincide. For domestic manufacturers, the World Artificial Intelligence Conference (WAIC) takes place in July, prompting companies to release models and showcase technological advancements around this time. For overseas manufacturers, several leading companies also choose to update their models during this period, aiming to gain a competitive edge in the developer ecosystem and the broader commercial landscape.

Moreover, the competitive landscape of the industry is evolving. Having a powerful model is no longer the sole objective; competition has expanded to how models are utilized and how they generate revenue.

This is also why the debate between open-source and closed-source models intensified again in July. For a long time, there has been a performance gap between open-source and closed-source models. But as some open-source models have entered the top tier, a more pressing question has emerged: when high-performance models are available for free, could model companies’ API usage and subscription revenues be diverted?

Supporters of open source argue that open weights lower technical barriers, enabling more companies and developers to participate in AI innovation. Open source means technology providers proactively cede some benefits to foster ecosystem growth; in contrast, closed-source proponents worry their users may be drawn away by open models. Companies like Anthropic believe that as model capabilities improve, releasing weights could pose security risks and require stricter governance.

Behind this debate lie the differing business interests of various companies. For infrastructure firms like NVIDIA, open models mean greater deployment demand and a larger compute market. For model companies like OpenAI and Anthropic, closed-source models help sustain revenue from API calls and subscriptions. However, it’s also important to recognize the other side: while open sourcing does increase deployment demand, improvements in parameter efficiency may reduce compute consumption per task, meaning the growth in the compute market may not directly correlate with the degree of model openness. For domestic companies, opening model weights is not merely a technical choice—it’s also a strategic move to gain access to the developer ecosystem, cloud services, and future commercialization opportunities.

After July, the competition among large models will continue. The developer community generally expects new models such as DeepSeek V4, Fable 5.1, and GPT-6 to be released陆续 in August.

The gap in model capabilities is narrowing, and it’s becoming increasingly difficult to establish a long-term advantage solely by releasing a stronger model. The key metrics to watch next are: how far the official release of DeepSeek V4 can push the upper limit of open-source models, whether API prices will continue to decline, whether agents can establish stable paid use cases, and whether open-source models can be adopted for private deployment by more enterprises. The answers to these questions will better reveal what this round of competition has truly changed than leaderboard rankings.

*Cover image sourced from Pexels.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.