FreeToken: open-source engine runs 290B+ MoE models locally on consumer hardware
tom_doerr · x · 2026-09-30
FreeToken (FlashML-org, 14k GitHub stars) is an open-source edge-native MoE serving engine that runs 290B+ open-weight frontier models on gaming PCs at interactive speeds. It treats GPUs, CPUs, host memory, and interconnects as one unified elastic inference platform, using bandwidth-adaptive CPU–GPU co-execution (q policy) and full-layer double buffering. Ships with installer, benchmarks, and an accompanying paper.
More from Infra
- Cerebras Stock Crashes Right Into Massive Lockup Expiration — firstadopter · 2026-09-30
- 5 ways to cut LLM costs without changing models: optimize tokens, caching and calls — goyalshaliniuk · 2026-09-30
- DeepSeek Now Training on Huawei Ascend 950 Chips — WebAssemblyMan · 2026-09-30
- EnerTune at SOSP'26 cuts LLM serving energy 1.4-2.3x vs SOTA systems — tianyin_xu · 2026-09-30
- SOSP'26 paper proposes energy-conscious GPU sharing for inference serving, beyond utilization — tianyin_xu · 2026-09-30
- SOSP'26 papers: AgileLog for isolated AI-agent data streams, CXL-LSM for CXL shared memory — tianyin_xu · 2026-09-30