FreeToken Lets 8GB GPUs Run 35B MoE Models On-Device

UC Berkeley and UT researchers released FreeToken, an inference engine that dynamically adapts GPU caches and bandwidth for MoE models, enabling an 8GB laptop GPU to run a 35B model at 39.3 tokens per second.

2026-08-29 ~ 2026-08-29 · 2 related posts

Full story(2 episodes)→