Open-Source Engine 'Waste' Runs Kimi K3 on Just 29GB of RAM
galapag0 · reddit · 2026-08-01
A developer shared an open-source project called Waste (Weight-Aware Streaming Tensor Engine) on Reddit. This engine is designed to run massive models on consumer-grade hardware.
Using this engine, users can run the Kimi K3 model with only 29 GB of RAM, achieving an inference speed of 0.50 tok/s. This provides a highly practical solution for deploying large MoE models on devices with limited VRAM or system memory.
Related event: Open-Source WASTE Engine Runs Trillion-Parameter Kimi K3 on Laptops(3 posts)→
More from Infra
- Cut Your $200/Month AI Bill: 10 Steps to Local AI Deployment — jason_mayes · 2026-08-01
- Local MXFP4 Testing Reveals Inconsistent Quality Among OpenRouter DS4 Flash Providers — antirez · 2026-08-01
- DeepSeek V4 Flash local benchmark nearly matches top frontier models from 5 months ago — joorklee · 2026-08-01
- New Method Pre-routes MoE Layers to Optimize I/O for Edge Streaming — dai_app · 2026-08-01
- 5TB of Data Stored on a Tiny Glass Slab Marks Microscopic Storage Breakthrough — TansuYegen · 2026-08-01
- antirez Enables Lossless MXFP4 Local Inference for DeepSeek v4 Flash — antirez · 2026-08-01