律动BlockBeats
律动BlockBeats|Jul 31, 2026 09:52
[Inference Engine WASTE Open-Sourced: Running Kimi K3 Fully on a 64GB MacBook] According to monitoring by Beating from 动察, edge database company SQLite AI has open-sourced the inference engine WASTE. It enables the full Kimi K3 model, retaining all layers and experts, to run on a 64GB MacBook Pro. The converted model takes up approximately 1TB, with a speed of 0.49 to 0.54 tokens per second. WASTE keeps about 27GB of the model backbone in memory, while over 80,000 experts are stored on the built-in SSD. For each token generated, only the currently invoked expert is read from the hard drive. This version does not involve distillation, pruning, or expert removal. However, expert weights have been re-quantized to 3-bit, while the model backbone uses 4-bit and 8-bit quantization. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads