Swiftlet reduces the 80B Qwen model to run on a Mac with a peak memory usage of 4.3GB

iconKuCoinFlash
Share
AI summary iconSummary
A new project announcement from MetaEra reveals that the Swiftlet project has optimized Qwen3-Next and other MoE models for Macs using Swift and Metal. By loading only a small dense core into memory and streaming expert weights from SSD, the 4-bit Qwen3.6-35B-A3B model requires just 18GB of disk space and 2.6GB of peak memory on an M5 Mac, achieving decoding speeds of 7–11 tokens per second. This on-chain update highlights efficient deployment of large models on consumer hardware.
ME AI message: The Swiftlet project leverages Swift + Metal to enable extreme optimization of MoE models like Qwen3-Next on Mac—keeping only a small dense core resident while streaming expert weights from SSD on demand. On an M5 Mac, the 4-bit version of Qwen3.6-35B-A3B occupies just 18GB of disk space, with a peak memory usage of 2.6GB and a decoding speed of 7–11 tok/s. (Source: MLion)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.