Qwen3.8-Flash-Next hits 68.3 tok/s on a single RTX 5090 with new FreeToken inference stack

matei_zaharia · x · 2026-09-05

UC Berkeley Sky Lab's Shuo demonstrated FreeToken, a local inference setup running Qwen3.8-Flash-Next on a single RTX 5090 at 68.3 tok/s — no extreme quantization, no speculative decoding. It uses a GB300-validated NVFP4 production checkpoint from RadixArk, needs only 63GB host RAM (less than 1-bit quants), and keeps the 51GB n-gram table on NVMe at 0.5% throughput cost. A demo shows generating a playable Minecraft-style world from one prompt and fixing a real FreeToken bug in 10.5 minutes. Matei Zaharia highlighted it as a sign powerful local AI is coming.

Original post →

More from Infra

Infra channel →