AI-generated summary: In July 2026, Moonshot AI open-sourced the Kimi K3 model and unveiled its list of 401 contributors for the first time. The company’s latest valuation reached $31.5 billion, with a team of over 300 members, an average age under 30, and each member contributing an average of $700 million in valuation. The core team includes three co-founders: algorithm lead Zhou Xinyu, architect Wu Yuxin, and CTO Zhang Yutao, along with key骨干 from top institutions such as Tsinghua University and CMU. The team has made multiple innovations in foundational technologies including the MoBA hybrid block attention mechanism, the Mooncake decoupled inference architecture, and the Muon optimizer, with related achievements recognized by honors such as the Best Paper Award at FAST 2025, a premier storage conference. This young team of enthusiasts is known for its flat management structure and has pushed engineering efficiency to the extreme under the principles of “long, fast, and lean,” establishing itself as a major force in China’s large model landscape.Article author and source: Leiphone
On July 27, Moonshot AI demonstrated its full commitment by open-sourcing the full-strength Kimi K3 model, a 2.8T parameter system. In the final section of their technical report, they revealed for the first time the official list of contributors behind Kimi K3—totaling 401 individuals—marking a comprehensive unveiling of Moonshot AI’s talented team.
From an initial startup team valued at just hundreds of millions of dollars, Moonshot has become the undisputed pride of China’s large model ecosystem—a elite team now capable of competing head-to-head with top-tier models like Claude 3.5 Sonnet and GPT-5.6 Sol on the global stage.
With the release of K3, Moonshot's latest valuation has surged to $31.5 billion, approximately RMB 210 billion, supported by an extremely young team.
According to publicly available information, Moonshot AI has over 300 employees, with an average age under 30, and 80% of them are I-type personalities. If calculated proportionally, each employee essentially supports a valuation of approximately 700 million RMB.
The organizational structure of Moonshot is highly flat. Moonshot has five founders, each directly managing 40 to 50 people. Internally, it is known for “no job titles, no KPIs, no departmental silos,” and employees can challenge their managers at any time. One detail: during a reflection meeting in early 2025, an employee proposed an aggressive suggestion; after listening, Yang Zhilin went even further, abandoning K1 entirely to focus on K2.
This eclectic spirit is also reflected in the company’s “soul.” As a rock enthusiast, Yang Zhilin has named meeting rooms after legendary bands like the Rolling Stones, Nirvana, and Radiohead, and even named Kimi’s membership tiers using musical terms such as Adagio and Allegro.
The Dark Side of the Moon believes that, in today’s world where computing power and data are increasingly homogenized, “taste” is the core driver of differentiation. This taste is not only reflected in the refinement of models but also deeply evident in their discerning and aspirational approach to talent.
Moonshot prefers hiring CEOs and strongly favors individuals with an "entrepreneurial aura" or "gambler's DNA." Public records show that at least 50 people within the company have founded or joined startups. In the past year, over 100 new hires at Moonshot joined through internal referrals.
Today, we have thoroughly researched and compiled all publicly verifiable information regarding the K3 contributors, as listed. The core background and formidable expertise of this elite team of technologists are fully revealed.
01 Kimi Technical System Foundational Pillar
Zhou Xinyu:
Zhou Xinyu is one of the core co-founders of Moonshot AI and also leads the algorithms team.
Zhou Xinyu and Yang Zhilin knew each other well; an interesting detail is that Zhou Xinyu was once the guitarist for Tsinghua University’s Splay band, with Yang Zhilin as the drummer.
Before joining Moonshot, Zhou Xinyu worked at Megvii on algorithm mass production and mobile model research. He closely collaborated with Zhang Xiangyu, then head of fundamental research at Megvii and now Chief Scientist at Jiepao星辰, and was a core author who personally wrote the lightweight network landmark paper ShuffleNet.
Zhou Xinyu specializes in the mass production of large model algorithms, excelling at efficiently and cost-effectively transforming cutting-edge laboratory algorithms into large-scale production systems. He is also skilled in hybrid architecture optimization; in a technical AMA on the Reddit community, he was the first to analyze the comparative advantages of combining the Kimi K3 KDA architecture with NoPE MLA, demonstrating that this architecture requires less computational power and operates faster in both pre-training and reinforcement learning.
Wu Yuxin:
Wu Yuxin is one of the co-founders of Moonshot AI. He is also a leading expert and renowned architect in the field of artificial intelligence, with deep expertise in the underlying engineering architecture and infrastructure of deep learning, computer vision (CV), and large language models (LLMs).
Wu Yuxin and Yang Zhilin share the same educational background, both earning their undergraduate degrees from Tsinghua University’s Department of Computer Science before pursuing further studies at Carnegie Mellon University (CMU) in the United States. He previously worked as a research engineer at Meta AI Research Lab (FAIR), where he led and contributed to several landmark technologies in computer vision.
Before co-founding Moonshot with Yang Zhilin, Wu Yuxin had already produced two industry-standard breakthroughs in deep learning and computer vision:
One of the primary creators and developers of Detectron2, this project directly reshaped the global computer vision research ecosystem, becoming the "universal foundation" for thousands of vision large models and object detection projects worldwide.
More innovatively, in 2018, he co-authored a landmark paper on Group Normalization with renowned scholar Kaiming He, earning a Best Paper Honorable Mention at ECCV 2018.
This classic paper directly overcomes the critical bottleneck of traditional Batch Normalization in small-sample training. The ability to “rewrite the global rules of AI training” with just a few lines of low-level innovation has given Wu Yuxin tremendous technical influence in the open-source community.
On the dark side of the moon, he was the one laying the foundation. He deeply contributed to optimizing large model architectures and training efficiency, and even attended the global PyTorch developer conference to thoroughly explain Kimi’s foundational contributions to the global open-source infrastructure.
Wu Yuxin’s expertise perfectly completes the final piece needed for Kimi to push the limits of next-generation multimodal long-context processing. He specializes in optimizing high-performance deep learning systems, solving how to make large models run more stably, compute faster, and use less memory—even with trillions of parameters—forming the technical backbone that ensures Kimi delivers a seamless experience under long-context and high-concurrency conditions.
As one of the core authors of MoCo v1, a landmark in deep learning, he has enabled Kimi not only to possess the strongest brain in the industry (natural language understanding) but also to be developing eyes capable of deeply understanding the physical world (multimodal vision).
Zhang Yutao:
Zhang Yutao is one of the co-founders of Moonshot AI and also serves as its CTO. He and Yang Zhilin are fellow alumni from Tsinghua University, and Tang Jie, founder of Zhipu AI, is his mentor.
In 2016, Zhang Yutao co-founded CircuSmart with Yang Zhilin, and later co-founded Moonshot AI with him. If Yang Zhilin is the “navigator” of Moonshot AI, then Zhang Yutao is the chief engineer responsible for fine-tuning the propulsion system of the “intelligence” vessel to its peak performance. Kimi’s groundbreaking advancements in long-text processing, complex reasoning, and agent task execution owe much to Zhang Yutao.
Before joining Moonshot AI, he had already delivered several industry-leading breakthroughs in the fields of knowledge graphs and heterogeneous data fusion.
One of his most well-known technical contributions was leading the development of AMiner, a big data analytics platform for academia, as the technical lead. The core heterogeneous data fusion technology behind this platform, which now serves millions of scholars worldwide, includes foundational capabilities that Zhang Yutao helped build during his Ph.D. studies at Tsinghua University. This deep understanding of “how massive amounts of knowledge can be absorbed, understood, and retrieved by machines” directly forms the technical foundation of Kimi’s long-text processing capabilities.
In terms of academic contributions, he has published multiple highly cited papers at top computer science conferences such as KDD and CIKM. He excels at enabling AI with "logical reasoning" through knowledge graphs and, during the era of recurrent intelligence, led the development of the "Thousand Recurrent Zero-Shot AI Platform," achieving large-scale deployment of industry-scale large models in enterprise service scenarios.
Zhang Yutao’s expertise precisely anticipated every milestone in AI’s transition from “able to converse” to “able to perform tasks.” He focuses on advancing long-context window technology and dynamic network architectures, addressing how to keep AI “alert” amid vast amounts of information.
At the latest technology summit in July 2026, he clearly outlined the three phases of industry engineering: 2023 was the Prompt era focused on “how to ask,” 2024 evolved into the Context era centered on “providing sufficient information,” and by 2026, large models have officially entered a new phase of Harness engineering. Today, he is leading his team to tackle the ultimate challenge of “how to build execution environments where AI can autonomously close the loop and fix its own issues.”
Xu Xinran:
Xu Xiran is the Vice President of Engineering and Head of AI Infrastructure at Moonshot AI.
He has full-stack foundational development experience and previously led the development of the MegEngine (SkyNet) deep learning framework and a billion-scale vector retrieval system at Megvii. Currently, he is primarily responsible for the overall architecture design of Moonshot AI’s large model training, inference, and infrastructure.
On the large model inference side, for Kimi’s long-text business scenario, he and his team jointly designed the separable inference architecture Mooncake, centered on KV Cache. This achievement, which completely separates the prefill and decode phases and alleviates memory and I/O pressure from million-token-long contexts, won the Best Paper Award at the storage conference FAST 2025.
At the same time, as one of the core authors of the MoBA (Mixture of Block Attention) mechanism, he significantly reduced costs and improved efficiency for Kimi at both training and inference stages through block-level attention partitioning.
On the model training side, he contributed to the development of the open-source project Moonlight and its underlying optimizer, Muon, replacing the traditional AdamW with matrix orthogonalization to reduce computational costs in distributed training of trillion-parameter models.
Based on his practical experience with distributed storage and inference under large-scale long-text workloads, he was invited to speak at the GOSIM China Global Open Source Conference. His expertise has always centered on full-stack large model architectures, aiming to enable high-intelligence long-text inference with reduced FLOPs overhead and lower KV Cache consumption.
He Weiran:
As the lead of the Moonshot Reasoning System, he is the primary author of the best paper stored at FAST 2025 and a core technical contributor to the engineering implementation of Kimi’s long-context reasoning architecture. From Kimi k1.5, Kimi-VL, and K2 to Kimi Linear and Kimi-Audio, and further to core architecture papers such as MoBA, Muon, and AttnRes, his name appears across all six generations of technological advancements—all centered on one core principle: trading system design for efficiency to make large models truly affordable.
Before joining Moonshot AI, He Weiran had spent many years specializing in high-performance computing. He graduated from Tsinghua University’s Department of Computer Science and initially focused on chip routing and low-level performance optimization. During his time at Megvii, he led efforts in developing proprietary AI chips and low-bit-width neural networks, working alongside Zhou Xinyu, who would later become a co-founder of Moonshot AI—Zhou Xinyu is the original inventor of the classic OCR text detection method, EAST.
After joining Moonshot AI in 2023, he applied his years of systems engineering expertise to the inference services of trillion-parameter large models.
His most representative achievement is the KVCache-separated inference architecture, Mooncake. This system breaks away from traditional inference service architectures by isolating and centrally managing the memory caches that consume the most GPU memory. Through global cache reuse and fine-grained routing, it increases Kimi’s request capacity by over 75% without increasing computational resources. At the 2024 QCon conference, he revealed that, leveraging this architecture and continuous engineering optimizations, the unit cost of Kimi’s inference cluster decreased by more than 20 times over the course of a year.
In February 2025, the Mooncake-related paper won the Eric R. Riedl Best Paper Award at the FAST conference—one of the highest honors in the field of storage systems. Even more indicative of its engineering value, this architecture was already stably running on thousands of nodes in Kimi’s production environment prior to the official award, processing tens of billions of user tokens daily.
Although the K3 technical report does not disclose specific individual responsibilities, several core breakthroughs build upon his team’s technical lineage: to address the challenge of traditional prefix cache invalidation in the KDA+MLA hybrid architecture, the team contributed an official adaptation implementation to the vLLM community; for million-token long requests, the cluster employs a cache-aware routing and scheduling strategy to maximize cache reuse efficiency. Ultimately, K3 delivered results with inference costs at just 38% of Claude Fable 5’s, with KVCache optimization, prefix caching, and cost control being core pillars of the Mooncake system.
From chip routing to trillion-parameter model inference, He Weiran’s technical path has always been clear: relentlessly optimizing system efficiency to achieve lower computational costs through engineering design. While top algorithm researchers raised Kimi’s capability ceiling, he and his team have made it their mission to bring this flagship-level performance down to earth—making it affordable and reliable for ordinary users.
02 Mid-level Offensive Team · On-the-Ground Key Personnel
Du Yulun:
Du Yulun is responsible for pretraining and scaling at Moonshot.
He graduated from the University of Illinois Urbana-Champaign and Carnegie Mellon University with a focus on artificial intelligence. He later joined Cycle Intelligence, serving as a researcher and head of the AI Lab, and is a core technical pioneer at Moonshot AI.
As the lead responsible for model base pre-training and scaling, his technical contributions fully span the architecture design of all Kimi core flagship base models and the entire suite of multimodal models, including VL and Audio.
He is a co-author of the papers "Attention Residuals," "MoBA (Mixture of Block Attention)," and "Muon is Scalable for LLM Training," which focus on breakthroughs in underlying technology. His work centers on alleviating signal degradation in ultra-deep networks through block-level attention partitioning and training flow reconstruction, significantly improving the distributed training efficiency of trillion-parameter models.
At the open-source engineering level, he is an active core contributor to Moonshot AI’s official open-source projects. Code collaboration records show that Du Yulun frequently collaborates with algorithm expert Su Jianlin and inference system lead He Weiran on co-development and code reviews, together forming the underlying technical闭环 that supports Kimi’s efficient processing of million-token-length texts and full-modal data flows.
Zhang Yu:
Zhang Yu is a core architect at Moonshot AI, specializing in the design of efficient, scalable, and hardware-friendly architectures for large language models. As a pioneer in non-Transformer and linear attention approaches, he is dedicated to overcoming the engineering challenge where longer texts result in slower performance and prohibitively high costs.
In both academic and engineering practice, Zhang Yu has achieved significant results. He is not only a co-first author of the paper "Attention Residuals" on residual connections, but also the lead developer of the efficient model architecture at Moonshot AI.
The Kimi Linear model, which he led as an open-source project, was the first to successfully apply linear attention mechanisms to ultra-long context tasks, achieving a revolutionary leap in efficiency. Moreover, his post titled "The Journey of Developing Kimi Linear," shared on technical communities like Zhihu, remains a must-read guide for countless large model data engineers.
From breaking through the fundamentals of implementing a 7B model to helping a trillion-parameter model enter the global top tier in a single battle, Zhang Yu effortlessly navigates between academia and industry, leaving behind numerous robust code contributions in top academic conferences and open-source communities.
This series of breakthroughs stems from his deep expertise in core technologies: in efficient model architecture design, he focused on optimizing communication and memory bottlenecks during model depth expansion; in training dynamics and efficiency improvement, he fundamentally resolved the issues of gradient dilution and hidden state explosion caused by PreNorm in large models by reconstructing the underlying residual flow.
Lai Guokun:
Lai Guokun is a dual alumnus of Tsinghua University and Carnegie Mellon University, like Yang Zhilin.
He is one of the researchers with the most publications on the Kimi team, with over ten papers. He is not only one of the representative figures of Moonshot’s “Genius Club” and a core contributor to Kimi K3, but also a legendary figure in academia with “defining” achievements.
Before joining Moonshot, Lai Guokun had already produced two industry-standard breakthroughs in the field of NLP (Natural Language Processing).
One of them created RACE, one of the world’s largest machine reading comprehension datasets, which trains AI using English exam questions from Chinese students and has become the global standard for machine reading comprehension.
More astonishingly, his undergraduate paper on recommendation algorithms, "Explicit Factor Models for Explainable Recommendation," received the SIGIR Test of Time Award in 2024—ten years after its publication. This is one of the highest honors in information retrieval, awarded exclusively to papers published a decade earlier that have stood the test of time and demonstrated lasting impact on the industry.
People with the ability to define the technological roadmap for the next decade are not uncommon at Moonshot, even at the undergraduate level.
Lai Guokun’s area of expertise aligns closely with Kimi’s core capabilities, such as his proficiency in "machine reading comprehension," which addresses how to enable AI to understand lengthy documents and answer questions—this is precisely the foundational technology behind Kimi’s long-text capabilities.
Pre-trained large models were one of his key research focuses during his time at CMU, along with neural architecture search—studying how to enable AI to automatically design better network structures—and interpretable recommendation systems, which allow AI to not only recommend content but also explain why a particular item was recommended.
Guo Haiqing:
Haiqing Guo is an early core R&D member of the Moonshot AI Kimi team and a veteran of the team, having been fully involved in the development and iteration of multiple generations of Kimi’s core projects.
Publicly verifiable participating projects include the Kimi K1.5 Technical Report, the Kimi K2 Technical Report, the AttnRes Core Architecture Paper, and the Kimi K3 Technical Report. AttnRes is one of the two core architectural innovations of Kimi K3.
Based on the contribution trail, he has been involved throughout the entire iterative cycle of Kimi, from the K1.5 to the K3 flagship generations, participating in core R&D stages across multiple product generations and remaining a consistent key contributor to Kimi’s technical team. His specific role and detailed R&D responsibilities have not yet been disclosed through public channels.
Chen Guangyu:
Chen Guangyu is likely the youngest researcher at Moonshot AI, born in 2009 at just 17 years old—a teenager who once impressed Elon Musk into publicly exclaiming "Impressive" on X.
When Moonshot unveiled its groundbreaking 2.8-trillion-parameter flagship model, Kimi K3, he was already listed as a co-first author of "Attention Residuals," the foundational paper behind Kimi K3’s core architecture, alongside tech experts such as Su Jianlin.
The attention residuals mechanism he developed directly reimagined the standard residual connections that had dominated deep learning for a decade, introducing "learnable deep attention" to enable models to intelligently "look back." This fundamental redesign boosted the model’s scaling efficiency by 1.25 times, allowing Chinese companies to move faster at lower cost in the trillion-dollar battle for top-tier computing power. His journey to Moonshot AI is extraordinary: just a year ago, he didn’t even know what a Transformer was. He was discovered by a Silicon Valley startup after proposing the idea of a "third mechanical auxiliary hand for humans" during a hackathon. Seven weeks later, the Kimi team took notice and exceptionally integrated him into the company’s core research group. In his first week on the job, he broke all conventions by personally rewriting key components and won first place in Kimi’s 48-hour extreme hackathon.
Du Angang:
In Kimi’s technical report, Du Angang’s name consistently appears first. He is one of the core contributors to Moonshot’s multimodal architecture and one of the few researchers to have participated in all five generations. He has twice been listed as the first author on team publications, suggesting his primary contributions to Kimi K3 lie in the visual domain.
Before joining Moonshot, Du Angang had already established two significant foundations in the field of computer vision.
One area is deep expertise in underwater microscopic vision. I spent seven years earning my bachelor’s and master’s degrees at the Underwater Vision Laboratory of Ocean University of China, specializing in the highly challenging task of classifying plankton microscopic images—enabling AI to accurately identify the smallest life forms in the ocean, even within blurry imagery, scarce samples, and extremely subtle species differences.
In 2020, he published related research as the first author in Global OCEANS, solving the fine-grained classification challenge using second-order feature modeling—the only such achievement in the lab’s decade of research in this area to be led by a master’s student as first author.
More impactful was his work during his time at Megvii. The multi-person pose estimation research he contributed to was accepted by ECCV 2020, a top conference in computer vision, and has accumulated over 300 citations, becoming a key reference in the field of human pose estimation. His collaborators on this work included leading figures in computer vision such as Sun Jian and Zhang Xiangyu, as well as Zhou Xinyu, co-founder of Moonshot AI.
After joining Kimi, he has focused on visual encoder design, multimodal alignment, and visual agent technologies, particularly excelling in high-resolution image understanding, long-sequence video parsing, and multimodal reinforcement learning—core technical foundations of Kimi’s native multimodal capabilities. In addition, his expertise extends downward to large model training infrastructure—from optimizer stability to agent data synthesis—forming a full-stack capability spanning from foundational visual algorithms to model training and up to agent deployment.
In 2025, he published the open-source multimodal model Kimi-VL as the first author, introducing the proprietary MoonViT vision encoder that natively supports ultra-high-resolution inputs. Combined with 128K long-context capability, this enabled Kimi to comprehend multimodal long-form content—such as lengthy videos and documents—for the first time, establishing the technical foundation for Kimi’s full suite of multimodal capabilities.
In the K2 phase, he actively contributed to the development of the MuonClip optimizer, collaborating with the QK-Clip mechanism to address the industry’s critical challenge of training instability in large models, enabling a massive 15.5 trillion-token training run with “zero loss spikes” and significantly reducing computational costs and risks. By the K3 phase, he led the development of an intelligent agent-driven data synthesis pipeline, allowing the model to autonomously generate high-quality training data—providing essential data support that enabled K3 to rank among the world’s top tier in demanding tasks such as programming and intelligent agents.
For Kimi, his contributions span the entire evolution of multimodal capabilities.
Qu Bowen:
Bowen Qu (Brian Qu) is part of the model post-training team. He earned his bachelor’s degree from the School of Electronic Information and Communications at Huazhong University of Science and Technology, and his master’s degree from the School of Electronics at Peking University. His research focuses on multimodal large models, specifically covering multimodal reinforcement learning, vision-language models, AI agents, and AIGC content quality evaluation.
After joining Moonshot AI, Qu Bowen’s core work has focused on building multimodal capabilities for the Kimi series of models, deeply participating in the Kimi-K2.5 project and contributing to research in native multimodal reinforcement learning. Kimi-K2.5 is Moonshot AI’s native multimodal agent model, with a total parameter scale of 1T and 32B activated parameters.
Qu Bowen’s native multimodal reinforcement learning system departs from the conventional approach of first performing visual supervised fine-tuning and then initiating reinforcement learning. Instead, it starts with a purely text-based reasoning model without any visual supervised fine-tuning, directly acquiring visual reasoning capabilities through visual reinforcement learning, thereby avoiding interference from low-quality visual data on the model’s native language reasoning abilities.
Qu Bowen's research began in the area of AIGC content quality assessment, with related成果 published at international conferences such as ICME 2024 and CVPRW 2024. The paper published at ICME 2024 addressed the limitation of existing AI-generated image evaluation systems, which focused solely on visual aesthetics and clarity while neglecting the alignment between generated content and user prompts. The study proposed integrating text prompts into the image quality evaluation framework, establishing a comprehensive standard that simultaneously assesses visual quality and instruction following, thereby enhancing the comprehensiveness of AI-generated content evaluation and improving alignment with human perception.
Qu Bowen is also listed as a contributor to the multimodal versions of the Kimi series. As a young researcher, Qu Bowen’s research journey has progressed from building evaluation frameworks to developing model capabilities, and ultimately to advanced agent research—achievements closely aligned with industry needs and directly supporting the iteration and enhancement of production-grade models.
Lu Enzhe:
Among the more than 400 contributors to Kimi K3, Lu Enzhe is not listed at the top of the contributor roster, but his contributions are exceptionally precise. Every achievement targets the core pain points of large model deployment, making him a key architect of Kimi’s high-efficiency computing infrastructure.
Lu Enzhe’s research focus can be precisely summarized in three words: longer, faster, cheaper—his core goal is to enable large models to process longer texts at lower costs and higher speeds. For Kimi, he is responsible for end-to-end efficiency improvements from training to inference.
In February 2025, he published the MoBA hybrid block attention mechanism as first author, transferring the expert selection concept from MoE to the attention layer, enabling the model to compute only the most relevant context blocks when processing requests. Achieving comparable performance to full attention at 95% sparsity, it accelerated inference by 6.5x for million-token inputs and 16x for ten-million-token inputs, directly solidifying the foundational architecture of Kimi’s ultra-long context capability.
On the day of its release, the paper debuted alongside DeepSeek’s NSA paper, becoming the most talked-about technical collision event in China’s large model community that year; its engineering value is further demonstrated by the fact that, prior to being accepted as a Spotlight paper at NeurIPS 2025, this solution had already been stably handling long-context user requests in Kimi’s production environment for nearly a year.
In addition, from MoBA to Kimi Linear’s KDA attention, he relentlessly optimized hardware inference efficiency throughout the architecture design. On the training side, his distributed implementation of the Muon optimizer achieved optimal memory usage and highly efficient communication, substantially reducing the computational cost of training large models. Including technical reports on AttnRes, K1.5, K2, Kimi-VL, and other generations, his publications comprehensively cover the entire pipeline of efficient large models.
Within the team, he is a consistent partner with Qiu Jiezhong, one of the corresponding authors of the MoBA paper—one takes the lead in execution, while the other guides the direction, and they work together seamlessly.
Wang Yuzhi:
Wang Yuzhi is a researcher at Moonshot AI and one of the most frequently cited authors on the Kimi series of technical reports.
From K1.5 in January 2025 to Kimi K3 in July 2026, his name appears on nearly every technical document produced by the company, covering vision, audio, attention architectures, training optimizers, flagship models, and cutting-edge AI risk governance frameworks.
He earned his bachelor’s degree from Xi’an University of Electronic Science and Technology and his Ph.D. in Electronic Engineering from Tsinghua University. After joining Megvii Research Institute, he served as a research engineer; in 2024, he joined Moonshot AI as a researcher, where he remains today.
Before joining Moonshot, his research focused on computational photography and low-level vision. Representative achievements include: as first author, publishing a paper on mobile RAW image denoising at ECCV 2020, which remains one of the benchmark works in mobile imaging; as a member of the Megvii team, winning the global championship in the raw-RGB track of the NTIRE 2019 Real Image Denoising Challenge; and as a co-author contributing to the seminal OCR work EAST (CVPR 2017, cited over 2,500 times).
During the dark moon phase, he co-authored ten technical documents, including: K1.5, K2, K3, Kimi-VL, Kimi-Audio, MoBA, Muon, Kimi Linear, four papers on architecture and efficiency, and a cutting-edge AI risk management framework.
Notably, he is listed as an author on all four of the company’s architecture and efficiency papers—a breadth of coverage that is uncommon among the approximately 400 contributors. His specific responsibilities within K3 are not publicly disclosed.
Mao Shaoguang:
Mao Shao Guang is a core member of the Kimi Agent technology team, fully involved in the development of four flagship generations: K2, K2-Thinking, K2.5, and K3. He is also one of the few researchers in the industry to have experienced both paradigm shifts in the Agent industry firsthand.
Before joining Moonshot, he was among the earliest practitioners in China to implement agents.
Graduated with a master’s degree in computer science from Tsinghua University, ranked in the top 0.2% of his class and awarded the National Scholarship. He served as a research assistant at The Chinese University of Hong Kong and completed an internship at Microsoft Research Asia’s Speech Group. After graduation, he worked at Microsoft for nearly six years, advancing from a research engineer to a senior development engineer.
In early 2023, participated in the development of TaskMatrix.AI to enable ChatGPT to access millions of APIs; subsequently co-proposed the Solo Performance Prompting method, allowing a single model to assume multiple roles and simulate multi-agent collaboration, with the technology subsequently integrated into Microsoft 365 Word.
That was the era of the Agent's "framework"—all capabilities relied on external tools stitched together and guided by prompts, and he was both an eyewitness and a builder of this technological approach.
But Mao Shao Guang, who had personally walked this path, was the first to see the limitations of the old paradigm. In June 2025, he wrote on Zhihu: “This direction is becoming increasingly uninteresting, with a sense of accelerating decline.”
After joining Moonshot AI in February 2025, he fully shifted to an end-to-end approach centered on “Model as Agent.” Within just four months, he contributed to the delivery of Kimi’s first Agent product, Kimi Researcher:
Trained entirely through end-to-end reinforcement learning without any predefined workflows, it averages 23 steps of reasoning per task, autonomously scans 206 web pages, and generates comprehensive 10,000-word research reports with properly formatted citations, achieving a 26.9% score on the HLE benchmark—setting a new global record at the time. He is also one of the two named authors of the product’s official technical documentation.
He was deeply involved throughout the development of the four flagship generations from K2 to K3. The HLE metric rose steadily from 8.6% to 26.9%—this performance curve, nearly entirely driven by reinforcement learning, is the record of his new technological approach.
From extending large models with external APIs to enabling models to inherently learn to use tools and autonomously complete complex tasks, Mao Shoguang’s career over the past decade has precisely mapped the full arc of the Agent industry’s evolution—from “assembly and combination” to “native growth.” Having personally participated in the old paradigm and then forged a new path, his insights into the industry are grounded in the most solid, real-world achievements.
Ruoyu Qin:
Qin Ruoyu is a Ph.D. candidate in the MADSys Lab at Tsinghua University’s Department of Computer Science, advised by Associate Professor Zhang Mingxing. His research focuses on efficient machine learning systems, specifically large-scale distributed large language model inference services and reinforcement learning rollout systems.
Mooncake, developed by Qin Ruoyu in collaboration with the Moon Shadow team, is a decoupled inference architecture centered on KVCache. This architecture separates the prefill and decode phases to run on independent clusters, while integrating idle CPU, memory, and SSD resources within the GPU cluster into a distributed cache pool. This achievement received the Erik Riedel Best Paper Award at FAST 2025, the premier conference in storage research.
Mooncake is a production-grade service platform already in live operation, with the paper’s abstract explicitly stating, “Mooncake is the serving platform for Kimi.” Under real-world workloads, Mooncake increases Kimi’s request processing capacity by 75%; in long-context scenarios, effective request handling capacity improves by 59% to 498%. The system is currently deployed across thousands of nodes and processes over 100 billion tokens daily.
In 2024, Moonshot AI and Tsinghua University's MADSys Lab jointly open-sourced the Mooncake project (open-source repository: kvcache-ai/Mooncake), which has since garnered over 6,000 GitHub stars. Relevant data and information can be found in the arXiv paper (ID: 2407.00079), the official announcement from Tsinghua University's Department of Computer Science, and the project's open-source repository.
After Mooncake, Qin Ruoyu’s research continued to deepen in the direction of reasoning systems. As first author, he authored the Seer paper, selected for OSDI 2026, which addresses rollout performance bottlenecks in reinforcement learning through rollout segmentation, context-aware scheduling, and grouped speculative decoding, achieving up to a 2.04x end-to-end rollout throughput improvement and reducing tail latency by 72% to 94%. Another first-author paper, Prefill-as-a-Service, further proposes a next-generation reasoning architecture design enabling KVCache sharing across data centers.
Qin Ruoyu also appears on the contributor list for the K3 project. The product performance of the large model is supported by two key capabilities: model development determines the upper limit of the model’s capabilities, while systems engineering determines the availability and operational cost of the service under high-concurrency scenarios.
Competition in the current large model industry is gradually shifting from a focus on training capabilities to a contest in service engineering. As a Ph.D. student, Qin Ruoyu has produced research成果 at the level of best-paper awards at top conferences, primarily because his work directly addresses production scenarios and is deployed in real business systems.
Yuan Enming:
Core R&D member of the Moonshot AI Kimi team, with research spanning computational biology, recommendation systems, and large models, and having fully participated in the technical evolution of multiple generations of Kimi’s core models.
Early in his career, he conducted research in AI-driven drug discovery in Professor Jia Yang’s group at the Institute for Interdisciplinary Information Sciences, Tsinghua University. Between 2020 and 2021, his research on a framework for influenza drug repurposing was published on the bioRxiv preprint server and later in the journal Signal Transduction and Targeted Therapy. He also published findings related to nucleocapsid protein phase separation in Protein & Cell.
From 2022 to 2023, Yuan Enming worked at Huawei Noah's Ark Lab, focusing on research in recommendation systems, with results published at multiple international top-tier academic conferences. His paper as first author on multi-behavior sequence recommendation was accepted at SIGIR 2022, and his collaborative works were also accepted at CIKM 2022 and WWW 2023.
After joining Moonshot, Yuan Enming actively contributed to the research and development of multiple generations of Kimi’s core models and technical projects. Publicly verifiable credited projects include the Kimi K1.5 Reinforcement Learning Technical Report, MoBA Technical Research, Kimi-VL Technical Report, Kimi K2 Large Model Technical Report, Kimi Linear Technical Report, Kimi K2.5 Multimodal Model Technical Report, and Kimi K3 Flagship Large Model Technical Report.
His scope of involvement spans multiple core technical areas, including reinforcement learning training, multimodal capabilities, foundational model architecture, and long-context systems, fully covering the three-generation flagship model iteration cycle from Kimi K1.5 to K3, making him one of the team’s core R&D personnel with the broadest technical coverage.
Yuan Enming’s specific role and detailed research responsibilities at Moonshot AI have not yet been disclosed through public channels. Based on publicly available authorship records, his interdisciplinary research background supports his involvement across multiple technical domains of large models, making him a key ongoing contributor to the evolution of Kimi’s technology framework.
Xu Weixin:
Xu Weixin is currently a researcher at Moonshot AI, focusing on large model architectures, efficient training mechanisms, and attention mechanism optimization.
Xu Weixin played a deep role in the research and development of Kimi’s multiple generations of core models and underlying technologies, with publicly credited contributions spanning the entire product iteration cycle from K1.5 to K3. After the release of K3, Moonshot AI officially disclosed three core proprietary technologies—the Muon optimizer, Kimi Linear linear attention, and Attention Residuals. Xu Weixin is the only verified contributor whose name appears in all three corresponding technical papers, ranked within the top 15 in each, making him one of the key developers behind the performance enhancements of the K3 architecture.
In terms of specific achievements, he was listed as the eighth author on the Muon Optimizer paper released in February 2025; this optimizer achieves approximately twice the computational efficiency of AdamW, and he also contributed to the subsequent open-source Moonlight training implementation.
In the Kimi Linear technical paper released in October 2025, he ranked 15th; the proposed KDA mechanism can reduce KV Cache size by up to 75% and increase decoding throughput by up to 6x in million-token scenarios.
In the 2026 March paper on Attention Residuals (one of the two core architectural innovations of K3), he ranked fourth, being the first key contributor after the three co-first authors; the related technology has been deeply integrated into the Kimi K3 foundational model architecture.
In addition, his name appears on the contributor lists of the Kimi k1.5 Reinforcement Learning Technical Report, the Kimi-VL Technical Report, the Kimi K2 and K2.5 Technical Reports, and the latest Kimi K3 flagship model technical report, with involvement spanning multiple core R&D areas including foundational models, multimodal learning, and reinforcement learning. It should be clearly noted that his name does not appear in the author list of the MoBA paper, and there is no publicly verifiable R&D connection between the two.
In addition, he has contributed to multiple research projects in 3D vision, including OccDepth for 3D semantic scene completion, PVD-AL for active learning distillation, and ACCV 2024 implicit SDF reconstruction, with academic outcomes spanning subfields such as model compression, 3D reconstruction, and semantic scene understanding.
Based on the public signature trail, Xu Weixin’s research trajectory has evolved from edge model deployment and 3D visual perception to innovations in large model underlying architectures, with long-term dedication to efficient computing. His technical expertise spans a broad range, making him a core and continuous contributor to the iterative development of Moonshot’s underlying technology stack.
Du Chenzhuang:
Du Chenzhuang is a core researcher at Moonshot AI and a key member of the team responsible for reinforcement learning and expanding reasoning capabilities.
He earned his bachelor’s degree in Information Management and Information Systems from Beihang University, then pursued his Ph.D. at the Institute for Interdisciplinary Information Sciences at Tsinghua University, known as the "cradle" of large model talent.
During his time at Tsinghua, he studied under the renowned AI scholar Zhao Xing and joined the prestigious MARS Lab (Multimodal and Autonomous Reasoning Laboratory), where he began focusing on multimodal machine learning, reinforcement learning, and reinforcement mechanisms for large language models.
During his doctoral studies, he served as a lead author in developing the groundbreaking ChatDB architecture, which pioneered the concept of using databases as symbolic memory modules to enhance large models, causing a sensation across both academic and industrial AI circles at the time.
To gain early hands-on experience with industrial-scale systems, he interned at the Beijing Haihua Artificial Intelligence Institute while pursuing his Ph.D., deeply contributing to a rigorous research project that built a foundational system from scratch. This enabled him to accumulate substantial experience in multimodal learning and reinforcement learning, while developing a bottom-up perspective integrating systems and algorithms.
After graduation, he joined Moonshot AI and frequently appeared in technical reports for core models. He directly contributed to the influential paper "Kimi k1.5: Scaling Reinforcement Learning with LLMs," exploring the scaling laws of reinforcement learning, and played a key role as a core contributor in the technical reports of flagship models such as Kimi K2 and K2.5.
He also played a key role as a core co-author in the development of the Kimi multimodal vision large model, the "Kimi-VL Technical Report," focusing on enhancing the large model’s perception and multimodal deep reasoning capabilities for complex images, charts, and real-world visual information.
Du Chenzhuang's primary focus is on large-scale large language model (LLM) training engineering, specializing in resource scheduling, stable training, and extreme system optimization for trillion-parameter models (MoE architecture) on superclusters; as well as reinforcement learning and multimodal fusion, researching how models can transcend their intelligence limits through self-supervision and reinforcement learning in complex long-horizon reasoning tasks, and achieving efficient underlying integration of visual information with language models.
Yanru Chen:
Yanru Chen is a core researcher at Moonshot AI, focusing on the development and implementation of large model architectures, long-text technologies, compute optimization, and multimodal engineering.
As a co-author of the paper "Attention Residuals," he contributed to redesigning the flow of information in the depth dimension of large models, successfully addressing the industry pain points of low training efficiency and information flow disruption in ultra-deep networks.
In the core technology of long-context processing, he co-authored MoBA (Mixed Block Attention) and led and participated in the development of Kimi’s most critical underlying mechanisms. By splitting attention at the block level, he directly addressed the excessive computational cost of million-token-long contexts during pre-training and inference, forming the foundational core of Kimi’s extreme cost reduction and efficiency gains.
At the same time, he is a co-author of the paper "Muon is Scalable for LLM Training" and played a key role in developing the influential open-source project Moonlight and its underlying Muon optimizer, which challenges traditional AdamW through matrix orthogonalization, effectively reducing the computational cost of training trillion-parameter models in distributed environments.
During the flagship model iteration, he deeply integrated the entire R&D lifecycle of Kimi’s previous flagship models (K1.5, K2, K2.5, K3), resolving challenges in stability and resource allocation under large-scale reinforcement learning and trillion-parameter scaling.
In addition, during the technical breakthroughs of Kimi’s multimodal lines (Linear, VL, Audio), he actively explored innovative linear attention mechanisms beyond traditional Transformer architectures and led the engineering implementation of Kimi’s visual and audio multimodal technology reports, efficiently resolving efficiency bottlenecks in底层 fusion of multimodal data streams.
Liu Shaowei:
Liu Shaowei is a core member of the MADSys Lab at Tsinghua University's top-tier Computer Systems Laboratory and a leading expert on the core infrastructure (Infra) team at Moonshot.
When you marvel at Kimi’s possession of trillions of parameters yet its remarkably low API pricing, it’s thanks to the team’s relentless pursuit of “extreme quantization art.” They have long specialized in ultra-low-cost inference for trillion-parameter models, native quantization, and large-scale distributed computing topology design.
On Reddit LocalLLaMA, one of the world’s most hardcore AI open-source communities, he published a groundbreaking technical deep-dive article titled “Kimi infra team: Quantization is not a compromise, it's the next paradigm” under the identity of a Kimi official expert. This in-depth article revealed the highly sophisticated native INT4 quantization format underlying Kimi K2 and K2.5.
He is not only a core co-author of MoBA (Mixed Block Attention Mechanism) but also deeply involved in the distributed architecture development of the flagship open-source project Moonlight and its underlying Muon optimizer, significantly reducing the computational cost of training trillion-parameter models through matrix orthogonalization.
He is also one of the core initiators and code maintainers of KTransformers, a major open-source project from Moonshot AI, and personally broke through the computational barrier for hybrid inference of trillion-parameter MoE models across heterogeneous clusters.
Chen Guanduo:
Chen Guanduo is a core researcher on the Infra team at Moonshot AI. He possesses exceptional cross-disciplinary R&D capabilities in "systems + algorithms," with research focused intensely on large model infrastructure, efficient attention architecture, acceleration of large model fine-tuning, and large-scale storage and retrieval optimization.
He earned both his bachelor’s and master’s degrees from the Department of Computer Science at Fudan University, establishing a solid professional foundation in distributed systems and underlying architecture. He later pursued advanced studies at Nanyang Technological University and the Hong Kong University of Science and Technology, where he studied under renowned scholars Yuan Binhang and Luo Siqiang in the fields of high-performance computing and large model infrastructure, gaining cutting-edge academic insights and research experience.
Before joining Moonshot AI, he underwent deep immersion in the core foundational teams of leading internet companies. He previously served as a research intern at ByteDance’s database systems development team and Meituan’s infrastructure team, actively contributing to the development and architectural optimization of industrial-grade, high-concurrency, large-scale distributed systems.
During this period, as a core author, he directly delivered two major breakthroughs. In terms of underlying storage, he developed and published "Oasis," introducing an optimal disjoint segment learning range filter that directly resolves the indexing bottleneck in large-scale range queries within underlying databases.
At the application layer, he published "Gar," significantly improving the accuracy of large models in precisely understanding human complex intentions and converting them into SQL database queries (Text-to-SQL) through a "generation and ranking" approach.
Chen Guanduo officially joined Moonshot AI in March 2025. As a core member of the Infra team, he is not only a primary co-author of Moonshot’s flagship large models, Kimi K2 and Kimi K2.5, but has also delivered a series of technical breakthroughs in areas such as large model fine-tuning acceleration, efficient attention architectures, and even final network architecture reconstruction.
To address the issue of excessive GPU memory consumption during large model fine-tuning, Chen Guanduo led the publication of the paper "Ce-LoRA," introducing a computationally efficient LoRA fine-tuning method. This technology significantly reduces GPU memory usage and computational overhead during parameter fine-tuning of large models, directly boosting training throughput at the infrastructure level—earning it the nickname "GPU memory slimming technique" for large model fine-tuning.
To address the problem of quadratic complexity explosion in long-text computation, he contributed to the development and publication of the groundbreaking open-source architecture, Kimi Linear—a highly expressive and efficient linear attention architecture that significantly reduces computational complexity for long texts while preserving the model’s powerful representational capacity. It successfully cuts KV Cache (key-value cache) memory usage by up to 75% and achieves 6x higher decoding throughput on texts exceeding one million tokens, forming the core technical foundation for Kimi’s ongoing advancements in long-context capabilities.
He is also one of the co-authors of "Attention Residuals."
From the底层 storage filters to computationally efficient fine-tuning algorithms, and up to revolutionary next-generation Transformer network architectures, Chen Guanduo’s career trajectory perfectly aligns with the industry paradigm shift in large models—from merely scaling parameters to pursuing极致 computational efficiency and practical engineering deployment.
Closing remarks
The K3 development team has always been low-key and enigmatic; during our review of the materials, it was difficult to clearly identify the specific roles and complete backgrounds of all members. Yet, through this substantial list, the answers have become clear: Why Moonshot? Why are they capable of supporting a super valuation of 700 million per person?
This group of young professionals, with an average age under 30, are system architects who truly understand the economics of computational power. From award-winning storage scheduling research presented at top conferences to foundational innovations in reimagining attention mechanisms, they have pushed engineering efficiency—longevity, speed, and efficiency—to its extreme.
It is precisely this end-to-end integration, from top-level algorithms to underlying hardware, that gives Moonshot AI the core advantage to break the curse of “compute homogenization.”
Another interesting detail is that at the end of the full roster, Kimi K3 has also become a key member of the entire R&D team. Creator and created now stand side by side on the same list—the golden age of China’s large models has already arrived with thunderous force.
