FreeToken Brings Frontier-Scale MoE Models to Laptops and Desktops, From 35B to 753B Parameters
青稞AI · wechat · 2026-10-09
Researchers from UC Berkeley, MIT and others have open-sourced FreeToken (at flashml.ai), an inference serving system natively designed for edge devices. Instead of treating a PC as a small GPU, it models the whole machine as a unified, elastic platform, co-designed around dynamic agent workloads and heterogeneous hardware.
- Key mechanisms: dynamic expert residency, CPU–GPU co-execution, agent state reuse, and runtime memory management — no fixed offloading policy
- Supports 20+ MoE models and runs real coding/tool-use agents on hardware from 8GB-VRAM laptops to single workstation GPUs
- Reported capability: 35B models on ordinary laptops, 284B on high-end gaming desktops, and a 753B GLM model on a single workstation GPU
The first author, UC Berkeley EECS PhD student Shuo Yang, will present the system in a community talk on October 11.
More from Infra
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- DIY hybrid GPU/CPU/SSD rig cuts DeepSeek TTFT from 75s to 8.9s at 16K prefill — HankYeomans · 2026-10-10
- vLLM thread (4/5): locality-domain MoE sharding speeds up decode 1.2x — vllm_project · 2026-10-10