FreeToken: Run 284B frontier models on consumer GPUs at interactive speeds
solyarisoftware · x · 2026-08-22
Researchers from UC Berkeley and MIT released FreeToken, an open-source MoE inference engine designed to run frontier-scale models on consumer hardware. Benchmarks show 39 tok/s for Qwen3.6-35B on an RTX 4060 and 25 tok/s for DeepSeek-V4-Flash 284B on an RTX 5090, claiming significant speedups over llama.cpp and Ollama.
More from Infra
- Token Usage Explosion: From 10B/Year to 10B/Week on OpenAI — xeophon · 2026-08-22
- GitHub Repo with 132k Stars Tracks Free Tiers for Devs — Roger_M_Taylor · 2026-08-22
- Switching to ROCm 10 Fixes Dynamic VRAM NaN on AMD GPUs — Present-Guitar-3967 · 2026-08-22
- NVIDIA Adopts Diamond Cooling for New Platform, Domestic Supply Chain Sees Boom — 创业邦 · 2026-08-22
- US AI industry's economic model is cracking: subscriptions below compute cost, data centers stalled — Minimum_Name9115 · 2026-08-22
- Starcloud runs a language model on its satellite after training on orbiting Nvidia GPUs — emmanuelvivier · 2026-08-22