FreeToken Paper: Edge-Native MoE Serving with Bandwidth-Adaptive Execution
mark_k · x · 2026-08-23
FreeToken is an edge-native MoE serving system designed to treat personal machines as unified, elastic inference platforms. The system co-designs the full serving stack around two realities of local AI: agent workloads continuously change execution patterns, and edge hardware exposes heterogeneous resources.
Rather than using a fixed offloading strategy, FreeToken dynamically maps computation and model state to available resources. It supports over 20 MoE models and real coding/tool-using agents across hardware from 8GB laptop GPUs to single workstation GPUs. Key results include running 35B models on laptops, 284B models on gaming desktops, and the 753B GLM-5.2 on a single workstation GPU.
More from Infra
- ClickHouse tops $350M ARR; OpenAI usage up 10x in circular AI financing — iamKierraD · 2026-08-25
- OpenAI's in-house inference chip reportedly rivals GB300, NVIDIA impact seen as limited — ivan_bezdomny · 2026-08-25
- Lambda seeks input on model cards: add NVFP4 weights and base models? — TheZachMueller · 2026-08-25
- Perplexity releases research on Portable Computer on Spark — AravSrinivas · 2026-08-25
- Arav Srinivas: On-device models critical for sensitive docs with SOTA OCR — AravSrinivas · 2026-08-25
- Arav Srinivas: Agentic inference must move to local hardware — AravSrinivas · 2026-08-25