FreeToken Engine: Run 35B Models at 39.3 t/s on 8GB GPUs
rohanpaul_ai · x · 2026-08-29
Researchers from Berkeley and UT Austin released FreeToken, a new engine designed to efficiently run large open models on consumer hardware by optimizing expert caching.
- Performance: An 8GB gaming laptop can run a 35B model at 39.3 tokens per second, outpacing production Codex speeds.
- Key Innovation: Unlike static engines, FreeToken allows the GPU cache to dynamically follow the router's requests, reducing missed experts from 62% (in llama.cpp) to 16%.
- Workflow: It dynamically splits remaining experts between the GPU and system RAM based on real-time PCIe and memory speeds, ensuring neither side idles while waiting for the other.
Related event: FreeToken Lets 8GB GPUs Run 35B MoE Models On-Device(2 posts)→
More from Infra
- Nvidia strategy: Buy the open-source layer underneath — bindureddy · 2026-08-29
- Meta uses Daytona sandboxes for Muse Spark multimodal reasoning model — pzakin · 2026-08-29
- Analyst updates hyper/neocloud software capabilities map for enterprise AI — BenBajarin · 2026-08-29
- Amallo: Open Source Tool for One-Click Remote Access to Local Ollama — tehfonsi · 2026-08-29
- Prediction Market: 68% Chance of State Data Center Moratorium by Year-End — Polymarket · 2026-08-29
- Conifer SDK Open-Sourced: Unified Gateway with Exact Cost Tracking — ycombinator · 2026-08-29