深潮TechFlow
深潮TechFlow|Aug 01, 2026 00:55
[WASTE Inference Engine Enables Local Execution of Kimi K3 2.78T Model on a 64GB Laptop] Deep Tide TechFlow reports on August 1st that, according to a trending project on Hacker News, the open-source inference engine WASTE has successfully enabled local inference of the full Kimi K3 model on a MacBook Pro equipped with 64GB of memory. According to the introduction, WASTE is developed in C language and achieves minimal memory usage of approximately 29.05 GiB by keeping the model backbone resident in memory and streaming MoE expert weights from disk on demand. The inference speed is approximately 0.49 to 0.54 tokens per second. The full Kimi K3 model contains approximately 2.78 trillion parameters, with a model container size of about 982 GiB, which was previously considered challenging to run on consumer-grade hardware. WASTE circumvents GPU memory limitations through its streaming mechanism, enabling laptops with standard configurations to run the full model weights without relying on quantized or pruned versions. Currently, the project has garnered significant attention in the Hacker News community, and the related code has been open-sourced.
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads