
Even before July has ended, the AI video generation sector has already become scorching hot.
Between July 3 and July 23, just 20 days, five companies announced major funding rounds. Keling AI Sensetime raised $3 billion, Shengshu Technology raised $500 million, Aish科技 secured hundreds of millions in its C+ round, Zhixiang Future raised RMB 1.5 billion in its C round, and even a startup less than a year old, FlovaAI, secured $80 million in its seed round.
The combined funding of the first four companies that reached unicorn status exceeded RMB 26.2 billion. In other words, the amount of capital flowing into the sector in just July 2026 surpassed the total raised over the entire previous two years (2024–2025).
Why are they all raising funds in July? What exactly is happening? This can't be a coincidence.
Keling AI: BAT makes historic joint investment; ARR quadruples in one year
July 3, Keling AI Announced a $3 billion Series A funding cap, led by CPE Source Peak and Guofang Innovation, with 34 institutions participating—Tencent, Alibaba Cloud, and Baidu notably appear together on the investor list. LightSource Capital served as financial advisor.
The Series A funding cap of $3 billion corresponds to a post-money valuation of approximately $18 billion; the signed portion to date amounts to approximately RMB 19 billion.
Why would capital dare to quote this number? It takes Kuaishou’s Q1 2026 earnings report to instill confidence in the market.
Kuaishou's Q1 2026 financial report disclosed, Keling AI Quarterly revenue exceeded RMB 650 million, a year-over-year growth of over 300%. The annualized revenue run rate (ARR) surged from $100 million in March 2025 to nearly $500 million in March 2026, quadrupling year-over-year.
This figure is unmatched in the AI video space—aside from ByteDance’s Seedance, no other domestic company has reached this scale.
Keling's revenue is driven by two pillars: B2B enterprise customer API calls and P-tier paid membership subscriptions. As of June 2026, Keling AI Over 100 million global users and nearly 50,000 enterprise clients. Keling has deeply contributed to the virtual environments and visual effects for the popular Chinese historical drama "Taiping Nian" and supported the generation of hundreds of high-quality shots in the Hollywood series "The Dynasty of David."
In terms of product development, in February this year, Keling launched its 3.0 series models, enabling video generation of up to 15 seconds, shot-level control, subject consistency, and synchronized audio-visual output.
Two details are worth noting. First, the valuation has been lowered from the $20 billion target rumored in May, making the pricing more pragmatic and accelerating the transaction. Second, the announcement includes a put option clause—should Beijing Keling fail to complete an IPO by October 2031, investors have the right to demand a repurchase of their shares at principal plus 8% simple annual interest.
Split, fundraising, equity incentives, and repurchase clauses—this comprehensive strategy clearly indicates that Ke Ling is preparing for an independent IPO.
Shengshu Technology: $500 Million Funding Before IPO, Transitioning from Video Models to “World Models”
On July 6, Shengshu Technology announced a $500 million B+ round of funding.
But if you look at the timeline more broadly, the company raised three rounds in five months: an A+ round exceeding RMB 600 million in February (led by Zhongguancun Science City and Xinglian Capital), a B round of RMB 2 billion in April (led by Alibaba Cloud), and this most recent round of $500 million in July, bringing the total raised to over RMB 5 billion.
Five months, three rounds—this pace indicates one thing: capital is rushing to get on board. Why the urgency?
At the end of March this year, Shengshu Technology completed its shareholding reform—a standard step before going public. Market rumors suggest it could initiate its Hong Kong IPO process as early as this上半年. As a result, investors are rushing to get on board before the IPO.
Shengshu Technology's core product is Vidu Launched globally in July 2024, it has now evolved through three iterations. Vidu S1 version — Real-time interactive model supporting voice-controlled frame adjustment and unlimited generation duration (1 minute to 2 hours).
Throughout 2025, Shengshu Technology achieved more than a tenfold growth in users and revenue, generating over 400 million videos in total. Its enterprise client base is particularly impressive: in advertising and e-commerce, it serves JD.com, Alibaba 1688, Amazon, Meituan, Focus Media, Blue Focus, L'Oréal, and Anta; in film and animation, it works with Tencent Animation, Wenxue Group, CCTV Animation, iQIYI, and Mango TV; in gaming, it supports Lilith Games and 37 Interactive Entertainment.
Shengshu Technology is not content with just video generation; this year, it began expanding into embodied intelligence.
In April 2026, Shengshu Technology officially launched the general world action model, MotuBrain, utilizing the World Action Model (WAM) approach to unify perception, prediction, and action into a single model, enabling robots with “one brain for foresight” and “one brain for multiple capabilities.” The commercial version of MotuBrain also entered industrial-grade validation.
This Tsinghua-affiliated company, founded just three years ago, is rapidly evolving from a video model company into a general world model platform. After receiving $500 million, the funds will be primarily allocated to the development of the general world model.
This is the real story that capital truly values.
Aishi Technology: A three-year-old unicorn with 150 million users worldwide
On July 14, Aishi Technology's C+ round was led by Alibaba. Combined with the C round completed in March, which raised $300 million (led by CDH Investments, with participation from nearly 20 investors including China Ruyi and 37 Interactive Entertainment), the total C-round financing reached RMB 2.98 billion, pushing the post-money valuation beyond $2 billion.
Alibaba previously led the Series B round of Aishi in September 2025 with a $60 million investment and is now increasing its commitment in this C+ round. Ant Group also led A2 round investment of over RMB 100 million as early as April 2024.
Alibaba's ongoing investment in Aishi demonstrates its strategic focus on Aishi as a key player in the AI video space.
As of March 2026, Aishi Technology has surpassed 150 million global users, with 15 million monthly active users and an annual recurring revenue (ARR) of $40 million.
Interestingly, co-founder Xie Xuzhang revealed that the training cost for Aishi’s models at this level is approximately only 10% of that of its peers. This cost advantage stems from filtering high-quality data, reducing ineffective training iterations, and leveraging an optimized Diffusion Transformer architecture to improve resource utilization and lower per-training costs, all achieved through the reuse of engineering expertise.
In January this year, Aishi released PixVerse R1, the world's first general-purpose real-time world model supporting 1080P. Unlike traditional pre-recorded video generation, R1 enables real-time dynamic generation—users can input new commands during video playback, and the scene achieves natural, smooth transitions in lighting and camera angles within approximately 0.5 seconds.

Aishi’s approach differs from other players: while others compete on generation quality, Aishi competes on interaction capability. Just one week after the R1 release, China Ruyi announced a strategic partnership with Aishi and invested $14.2 million.
The latest Series C+ funding will be used specifically for video generation foundational models, real-world models, and global growth.
Zhixiang Future: A New Unicorn with a Unique Approach of “Model + Agent + Hardware”
Zhi Xiang Future is a "new unicorn" on this list—after securing its C-round funding of RMB 1.5 billion, its post-investment valuation exceeded US$1 billion. The founder, Mei Tao, is a Foreign Member of the Canadian Academy of Engineering and former Vice President of JD.com.
The company has completed three rounds of funding in the past three months, totaling over RMB 2.1 billion.
On July 23, Zhi Xiang Future announced a C-round financing of RMB 1.5 billion, led by the National Social Security Fund, Sichuan Industrial Revitalization Fund, and ICBC Investment, with participation from film and television industry investors such as Shanghai Film New Vision Fund and Huace Pictures, along with 18 other institutions—a combination of national long-term capital, local state-owned capital, and industry capital that is uncommon in the AI video sector.
Zhi Xiang launched the world's first open-access video generation DiT architecture model in May 2024 and unveiled the multimodal creation agent Vivago R1 at the recently concluded WAIC 2026, featuring unlimited-length video generation and editing.
Its products currently serve over 50 million users and 40,000 enterprise customers across more than 100 countries worldwide. Total revenue for 2025 exceeded RMB 100 million, and first-quarter revenue in 2026 has already surpassed that of the entire previous year. During the 2026 Chain Expo, Mei Tao stated that due to the high cost of large model pre-training, the company expects to achieve monthly break-even by 2029.
Mei Tao also spoke directly: “We’re not competing with ByteDance on underlying capabilities; we’re focused on commercial marketing and professional film collaboration.”
Zhi Xiang’s strategy is clear: avoid direct competition with major players and focus deeply on vertical use cases. It first built its user base with image models, then progressed toward video and world models.
FlovaAI: $80 million seed round, Guo Lie's third venture
Guo Lie, the founder of FlovaAI, is an unavoidable name in the history of China's mobile internet. In 2013, Guo Lie created FaceMe, followed by FaceU JiMeng. In 2018, his team was acquired by ByteDance for $300 million. Since then, he has been involved at ByteDance in incubating LightCam, PicWish . CapCut ——That's right, CapCut He was the one who helped create it.
In 2025, Guo Lie set out again, founding Shenzhen Yuzhou Technology, and launched Flova.ai in October. Sequoia Capital China, IDG Capital, and Yunji Capital have collectively invested over $80 million.
Even at the angel round, they offered $80 million—not betting on the product, but on Guo Lie himself. Every consumer-facing tool he has ever built has become a phenomenon.

Flova is an AI-native video creation agent platform. Users express their creative needs through conversation, and the agent handles the entire process—from scriptwriting and character design to storyboard creation, generation of images/video/audio, and assembly of the editing timeline.
Unlike the mainstream canvas layouts on the market, Flova has designed a “storyboard” workspace. Agents receive instructions in a chat window, their workflow is displayed on the left storyboard, and users can view results directly in the central preview area; if unsatisfied, they can double-click at any time to intervene and make edits.
Multiple early users have noted that Flova’s standout feature is its “memory”—even after creating dozens of episodes, it can still recall specific storyboard details from the first episode. This long-context memory capability is highly distinctive among similar products.
At this stage, Flova is adopting a zero-margin strategy. Guo Lie said, “During the development phase, we aim to maintain zero margins.” The team is not focused on short-term revenue but rather on the effective consumption ratio: what portion of token usage reflects genuine user value, and what portion represents wasteful spending due to product shortcomings.
Why are everyone crowded in July? Possible reasons
Now, back to the core question: Why did these five companies all conduct concentrated fundraising in July? When I pieced together all the information, I found that four reasons converged in resonance.
First, ByteDance's "siphoning effect" forces everyone to accelerate.
Since its release in February this year, ByteSeedance 2.0 has generated over RMB 1 billion in monthly revenue, achieved approximately 95% market penetration in the short video industry, and holds over 80% market share based on daily token consumption.
Volcano Engine has raised its MaaS revenue target for 2026 from RMB 1.5 billion to RMB 15 billion—a tenfold increase, nearly entirely driven by the Seedance model. Additionally, Seedance 2.1 is set to launch soon, with expected performance improvements of 20% and a low-end version priced at just 0.5 yuan per second.
While ByteDance is dominating the market with a dual strategy of superior model capabilities and price competition, other players risk losing their place at the table entirely if they don’t quickly secure funding, accumulate computing power, and accelerate iteration.
Everyone is looking for a long-term counterbalance to ByteDance—but each takes a different approach: KeLing pursues a platform-based, multimodal strategy, ShengShu focuses on a technical enthusiast approach, AiShi emphasizes real-time interactive differentiation, ZhiXiang concentrates on deep vertical applications, and FlovaAI follows an agent workflow strategy.
Capital is not investing in the same story; it's betting on different technological pathways and business models.
Second, the business model has been successfully validated, and investors see a clear exit path.
Last year, people were debating whether AI video could generate revenue; this year, the data is in: Keling’s ARR is nearly $500 million, Shengshu’s revenue has grown more than tenfold by 2025, Aishi’s ARR is $40 million, and Seedance has surpassed $1 billion in a single month.
Advertising, e-commerce, short dramas, gaming, and film and television—AI video monetization scenarios are being validated one by one. This visible cash flow has directly triggered concentrated capital investment.
Third, the narrative has evolved from "video generation" to "world models."
In the first half of this year, the competitive landscape for foundational large models became more concentrated, and venture capital began seeking the next direction, with multimodal generation and world models emerging as consensus pathways.
Note a trend: None of the four companies that secured massive funding describe themselves solely as working on "video generation." Keling is talking about an all-in-one multimodal creation workflow, Shengshu is focusing on "general world models" and embodied intelligence, Aishi is emphasizing "real-time world models" and interactive entertainment, and Zhixiang is highlighting "native multimodal world models" and causal reasoning.
Video generation is just the entry point; the true endgame is a world model—an AI system capable of understanding, reasoning, and constructing physical worlds. Investors aren’t looking at this year’s video generation market; they’re betting on which world model will emerge in three years.
Fourth, the Hong Kong stock IPO window has opened, and capital is racing to secure tickets.
In January this year, Zhipu, MiniMax After listing on the Hong Kong Stock Exchange, stock prices surged sharply; the amount raised through HK IPOs in Q1 broke HK$100 billion in just 79 days, setting a new historical record for the fastest pace, with over 350 companies waiting to list.
Shengshu has completed its share restructuring, with persistent rumors of an IPO; Ke Ling’s split financing combined with repurchase clauses is essentially a prelude to listing. The intensive fundraising in the primary market is largely in preparation for entry into the secondary market.
Finally, I’d say that after this July, the AI video sector has entered the final round, and the landscape will likely become much clearer: leading companies with substantial funding are racing toward IPOs, mid-tier companies are being revalued under the narrative of world models, and new entrants like Flova are attempting to bypass the model arms race by using agents to carve out a slice of the market at the application layer.
But behind the excitement lies a problem: money comes in quickly, but it’s burned even faster. Video generation is the most computationally intensive category of AI; while $26 billion sounds staggering, when spread across each company’s annual compute costs, it may not be enough.
This article is from the WeChat public account "IT Juzi" (ID: itjuzi521), author: Wu Meimei.
