FreeToken: Open-Source Engine Runs 290B+ MoE Models Locally on Gaming PCs, 13.8k Stars
solyarisoftware · x · 2026-09-26
FlashML-org released FreeToken, an open-source edge-native MoE serving engine (13.8k GitHub stars) that brings datacenter-scale model serving to desktops. It treats GPUs, CPUs, host memory, and interconnects as one elastic inference platform to run 290B+ open-weight frontier MoE models on consumer hardware at interactive speeds, using bandwidth-adaptive CPU–GPU co-execution (q policy), full-layer double-buffering and other edge optimizations. Ships with one-line install, benchmarks, docs, and a linked paper.
More from Infra
- NVIDIA at $5.4 trillion is now worth more than the entire UK or French stock market — iamfakhrealam · 2026-09-26
- DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests — teortaxesTex · 2026-09-26
- A Curated Paper List for Learning Distributed LLM Training and Inference — East-Muffin-6472 · 2026-09-26
- vLLM adds Elastic Expert Parallelism: grow/shrink MoE GPU pools under live traffic — PyTorch · 2026-09-26
- AI data center investors now favor real infrastructure over PowerPoint promises — TansuYegen · 2026-09-26
- Running Qwen 27B Q4_K_M on dual RTX 3060 with llama.cpp hits ~44-50 tok/s — jacek2023 · 2026-09-26