FreeToken Engine: Run 290B+ Frontier MoE Models Locally on 8GB Laptop GPUs

Saboo_Shubham_ · x · 2026-08-25

FreeToken is an edge-native Mixture-of-Experts (MoE) serving engine designed to run frontier-scale open-weight models on consumer hardware. It enables running frontier models on laptops with just 8GB VRAM, supports OpenAI and Anthropic API compatibility (e.g., Claude Code, Codex), and utilizes techniques like CPU-GPU co-execution and global LRU expert caching for optimization. The project is 100% open-source.

Related event: FreeToken Runs 290B MoE Models Locally on 8GB GPUs(2 posts)→

Original post →

More from Infra

Infra channel →