FreeToken: 4x Faster Decode, Enables 284B Models on Gaming Desktops
airesearch12 · x · 2026-08-22
FreeToken is an edge-native MoE serving system that achieves 3–4× faster decode and 6–30× faster prefill compared to Ollama. It utilizes bandwidth-adaptive CPU–GPU execution and semantic-aware caching across agent turns.
By treating personal machines as elastic inference platforms, FreeToken dynamically maps computation and model state. It enables running a 35B model on an 8GB laptop GPU, a 284B model on a gaming desktop, and the 753B GLM-5.2 on a single workstation GPU.
Related event: Berkeley and MIT Open-Source FreeToken to Run 284B Models on Consumer GPUs(4 posts)→
More from Infra
- Ox Alpha's 100T tokens/day giveaway costs $1-2M in electricity at full capacity — teortaxesTex · 2026-08-22
- Seeking unified AI gateway for OpenAI cost visibility — HurryOrganic · 2026-08-22
- Agents fail silently: $47k loop reveals monitoring gaps — alifgokce · 2026-08-22
- Anti-datacenter movement is humanity recognizing its successor — ZeroStateReflex · 2026-08-22
- Can llama.cpp share KV cache across multiple GPUs for parallel requests? — spaceman_ · 2026-08-22
- GLM-5.2 local inference: ubatch size significantly boosts MoE performance — fuzhongkai · 2026-08-22