SQLite AI open-sources the WASTE inference engine to run Kimi K3 on a 64GB MacBook

iconKuCoinFlash
Share
AI summary iconSummary
SQLite AI has open-sourced its WASTE inference engine, enabling the full Kimi K3 model to run on a 64GB MacBook Pro. The model, converted to 1TB, operates at 0.49–0.54 tokens per second. WASTE loads only the required experts from SSD, keeping the 27GB backbone in memory. Weights are re-quantized to 3-bit and 4–8-bit. This AI and crypto development represents a step forward in on-chain news for model efficiency.
ME AI News: According to monitoring by Beating, edge database company SQLite AI has open-sourced its inference engine, WASTE. It enables Kimi K3, which retains all layers and experts, to run on a 64GB MacBook Pro. The converted model occupies approximately 1TB and achieves a speed of 0.49 to 0.54 tokens per second. WASTE keeps about 27GB of the model backbone in memory, while over 80,000 experts are stored on the built-in SSD. Only the currently invoked experts are read from the disk for each generated token. This version does not involve distillation, pruning, or removal of experts; however, expert weights have been re-quantized to 3-bit, and the model backbone uses 4-bit and 8-bit quantization. (Source: MLion)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.