ME News reports that on July 29 (UTC+8), Moonshot AI launched Kimi Linear, a hybrid linear attention architecture that outperforms full attention mechanisms across short-context, long-context, and reinforcement learning scenarios for the first time. Its 3B activated parameter model significantly surpasses full MLA on all evaluation tasks, reduces KV cache usage by up to 75%, and achieves up to 6x higher decoding throughput with a 1M context length. Moonshot AI has open-sourced the KDA kernel, vLLM implementation, and model weights. 🔗 Read the original article via AI HOT · https://aihot.virxact.com/items/cms4t5zra01o1roa10b19hmbh (Source: AiHot)
Moonshot Launches Kimi Linear, a High-Performance Attention Architecture
KuCoinFlashShare
On July 29 (UTC+8), Moonshot launched Kimi Linear, a high-performance attention architecture. The hybrid linear model outperforms full attention in both short- and long-context scenarios and reinforcement learning. With 3 billion active parameters, it reduces KV cache usage by 75% and increases decoding throughput by 6x at a 1M context length. The KDA kernel, vLLM implementation, and model weights have been open-sourced. Traders assessing support and resistance levels may find this advancement enhances the risk-to-reward ratio in algorithmic strategies.
Source:Show original
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information.
Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.