FreeToken: 8GB Laptop Runs 35B Model at 39.3 t/s via Edge-Native MoE Serving

rohanpaul_ai · x · 2026-08-29

Researchers from Berkeley and UT Austin released FreeToken, an edge-native MoE serving system that optimizes local inference via bandwidth-adaptive execution. By dynamically mapping computation and state to heterogeneous resources, it enables a 35B model to run at 39.3 t/s on an 8GB laptop and even a 753B model on a single workstation GPU, significantly lowering the hardware barrier for local frontier AI.

Related event: FreeToken Lets 8GB GPUs Run 35B MoE Models On-Device(2 posts)→

Original post →

More from Infra

Infra channel →