FreeToken Enables 35B Models on 8GB Laptops via Unified CPU-GPU Memory
mark_k · x · 2026-08-23
FreeToken is a software breakthrough for local AI that intelligently coordinates GPU, CPU, and system RAM to dynamically move MoE experts where needed, eliminating the requirement for enough VRAM to hold an entire model.
Benchmarks show an ordinary laptop with an 8GB RTX 4060 running a 35B model at 39 tok/s, while a gaming PC with an RTX 5090 can interactively run a massive 284B DeepSeek model. This allows users to run powerful AI on existing hardware without API costs or cloud dependency.
More from Infra
- Passive Oculink setup unintentionally creates CUDA task scheduler — TinFoilHat_69 · 2026-08-29
- RX 9070 XT shows huge speed variations with ComfyUI — bosox62 · 2026-08-29
- Open Source Static Performance Model for LLM Inference — stanfordnlp · 2026-08-29
- Prediction: Closed frontier models to become downloadable by 2027 — imjustnewatai · 2026-08-29
- Achieving 181 tok/s on Qwen3.8 with 2x DGX Sparks via NVMe offloading — StartupTim · 2026-08-29
- Together AI processes 135B+ GLM-5.3 Flash tokens in 24 hours — togethercompute · 2026-08-29