FreeToken: 8GB Laptop Runs 35B Model at 39.3 t/s via Edge-Native MoE Serving
rohanpaul_ai · x · 2026-08-29
Researchers from Berkeley and UT Austin released FreeToken, an edge-native MoE serving system that optimizes local inference via bandwidth-adaptive execution. By dynamically mapping computation and state to heterogeneous resources, it enables a 35B model to run at 39.3 t/s on an 8GB laptop and even a 753B model on a single workstation GPU, significantly lowering the hardware barrier for local frontier AI.
Related event: FreeToken Lets 8GB GPUs Run 35B MoE Models On-Device(2 posts)→
More from Infra
- Lambda raises $1B debt to buy Nvidia chips for Microsoft — himanshustwts · 2026-08-29
- RX 9070 XT shows huge speed variations with ComfyUI — bosox62 · 2026-08-29
- Open Source Static Performance Model for LLM Inference — stanfordnlp · 2026-08-29
- Prediction: Closed frontier models to become downloadable by 2027 — imjustnewatai · 2026-08-29
- Achieving 181 tok/s on Qwen3.8 with 2x DGX Sparks via NVMe offloading — StartupTim · 2026-08-29
- Together AI processes 135B+ GLM-5.3 Flash tokens in 24 hours — togethercompute · 2026-08-29