Qwen3.8-Flash-Next hits 68.3 tok/s on a single RTX 5090 with new FreeToken inference stack
matei_zaharia · x · 2026-09-05
UC Berkeley Sky Lab's Shuo demonstrated FreeToken, a local inference setup running Qwen3.8-Flash-Next on a single RTX 5090 at 68.3 tok/s — no extreme quantization, no speculative decoding. It uses a GB300-validated NVFP4 production checkpoint from RadixArk, needs only 63GB host RAM (less than 1-bit quants), and keeps the 51GB n-gram table on NVMe at 0.5% throughput cost. A demo shows generating a playable Minecraft-style world from one prompt and fixing a real FreeToken bug in 10.5 minutes. Matei Zaharia highlighted it as a sign powerful local AI is coming.
More from Infra
- Bad Apple running on a Trainium chip: real kernel instruction trace executes every frame — daniel_c0deb0t · 2026-09-05
- Dell ships world's first production NVIDIA Vera Rubin NVL72 racks to CoreWeave — rohanpaul_ai · 2026-09-05
- Paraguay deploys 1,600 Starlink kits to reach 50,000+ students and teachers — elonmusk · 2026-09-05
- How ±0.1°C thermal control and electrostatic chucks make next-gen AI chips possible — bookwormengr · 2026-09-05
- How lithography wafer stages float: air bearings and maglev enable nanometer-scale positioning — bookwormengr · 2026-09-05
- FrankenTTS: pure-Rust Qwen3-TTS port runs voice cloning entirely in your browser — BLUECOW009 · 2026-09-05