Qwen3.8 Flash Next on one RTX 5090 hits 50 t/s decode via FreeToken expert caching

dir3ctly · reddit · 2026-09-20

A Reddit user shares benchmarks running Qwen3.8 Flash Next locally on a single RTX 5090 with FreeToken: 50 t/s token generation and 2300 t/s prefill, stable over long contexts.

Original post →

More from Infra

Infra channel →