FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
SteppenAxolotl · reddit · 2026-08-25
FreeToken is an edge-native MoE serving framework that achieves significant speedups through bandwidth-adaptive CPU-GPU execution and semantic-aware caching across agent turns.
Performance Claims:
- Qwen3.6 35B: 8GB RTX 4060 laptop @ 39 tok/s
- DeepSeek-V4-Flash 284B: RTX 5090 desktop @ 22-25 tok/s
- GLM-5.2 753B: RTX PRO 6000 workstation @ 15 tok/s
It claims 3–4× faster decode and 6–30× faster prefill compared to Ollama, enabling frontier models on consumer hardware without extreme quantization.
Related event: FreeToken: Open-Source Engine Runs 290B MoE Models Locally on 8GB GPUs(3 posts)→
More from Infra
- AMD showcases memory fabric design optimizing memory resources — BenBajarin · 2026-08-25
- AMD details MI455X GPU and Helios rack-scale system at Hot Chips — firstadopter · 2026-08-25
- Nvidia's Cost Advantage Remains Even if Competitor Chips Were Free — sinclairx · 2026-08-25
- Lium: A Vast/Runpod Competitor Offering Cheaper GPUs and No KYC — markjeffrey · 2026-08-25
- Vera Rubin cooling loop operates with just 10°C delta — beffjezos · 2026-08-25
- NVIDIA Rubin's power smoothing tech changes data center economics — BenBajarin · 2026-08-25