FreeToken Paper: Edge-Native MoE Serving with Bandwidth-Adaptive Execution

mark_k · x · 2026-08-23

FreeToken is an edge-native MoE serving system designed to treat personal machines as unified, elastic inference platforms. The system co-designs the full serving stack around two realities of local AI: agent workloads continuously change execution patterns, and edge hardware exposes heterogeneous resources.

Rather than using a fixed offloading strategy, FreeToken dynamically maps computation and model state to available resources. It supports over 20 MoE models and real coding/tool-using agents across hardware from 8GB laptop GPUs to single workstation GPUs. Key results include running 35B models on laptops, 284B models on gaming desktops, and the 753B GLM-5.2 on a single workstation GPU.

Related event: UC Berkeley and MIT Open-Source FreeToken: Frontier MoE Models on Consumer PCs(11 posts)→

Original post →

More from Infra

Infra channel →