SQLite AI open-sources inference engine WASTE for Kimi K3 on 64GB MacBook

icon MarsBit
Share
AI summary iconSummary
SQLite AI has open-sourced its inference engine WASTE, enabling the full Kimi K3 model to run on a 64GB MacBook Pro. This on-chain development represents a significant milestone in AI + crypto news, as the model operates at 0.49 to 0.54 tokens per second. The 1TB model retains the 27GB backbone in memory and loads experts from SSD as needed. Weights have been re-quantized to 3-bit, with the backbone using 4-bit and 8-bit precision.

According to Beating Monitor, edge database company SQLite AI has open-sourced its inference engine, WASTE. It enables the full-layer, full-expert Kimi K3 model to run on a 64GB MacBook Pro. The converted model occupies approximately 1TB of storage and achieves a speed of 0.49 to 0.54 tokens per second. WASTE keeps about 27GB of the model backbone in memory, while over 80,000 experts are stored on the built-in SSD. Only the currently invoked experts are read from the disk for each generated token. This version does not involve distillation, pruning, or removal of experts; however, expert weights have been re-quantized to 3-bit, while the model backbone uses 4-bit and 8-bit quantization.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.