FreeToken: run 290B+ frontier MoE models locally on your gaming PC
Saboo_Shubham_ · x · 2026-08-25
FreeToken is an edge-native MoE serving engine (4.8k stars, 422 forks on GitHub) built to run frontier-scale open-weight models on consumer hardware — claiming interactive-speed inference of 290B+ parameter MoE models on a gaming PC.
Key design points:
- Treats heterogeneous edge resources (GPU, CPU, host memory, interconnects) as one elastic inference platform
- Bandwidth-adaptive CPU–GPU co-execution with a q policy
- Full-layer double-buffered prefill streaming, global LRU expert caching, and the FTW fast weight format
- Semantic-aware caching
It speaks the Anthropic/OpenAI APIs so Claude Code and Codex plug straight in; fully open source with a paper and community channels. The thread also highlights the Awesome LLM Apps repo (134k+ stars) with 100+ AI agent, skills and RAG templates.
Related event: FreeToken Runs 290B MoE Models Locally on 8GB GPUs(2 posts)→
More from coding & agent
- Paper: Agent Memory Provenance Has a Budget Problem — agihouse_org · 2026-08-25
- Build a Multi-Agent GTM Intelligence System to Boost Sales — LightningAI · 2026-08-25
- Naming Methods Hurts Agents: Steps Drive Performance — rohanpaul_ai · 2026-08-25
- Codex Computer Use installs Diablo II on macOS via Wine — Dimillian · 2026-08-25
- Hyo: Turning Obsidian into an open-source Claude Agent OS — evielync · 2026-08-25
- Lovable hits 250 repos/sec peak using Code.Storage for AI coding infra — dhruv2038 · 2026-08-25