87GB Qwen3.8 Flash Next runs at 120-150 t/s on a single RTX 5090
CurieuxExplorer · x · 2026-10-01
A Japanese user tested the local inference tool Strata: Qwen3.8 Flash Next (87GB) runs at 120-150 t/s on a single RTX 5090 with 4K t/s prefill. Hooked into LibreChat with web search, it retrieves properly and has relaxed guardrails. The takeaway: Qwen-class models are now a desktop sport, undercutting the "you need a data center" era.
More from Infra
- Arduino argues sub-$900 embedded boards beat Mac minis for Physical AI agent economics — CatAstro_Piyush · 2026-10-01
- SemiAnalysis injects failures to rate GPU cluster renters — providers differ sharply on recovery — AccBalanced · 2026-10-01
- Redditor runs unattended DeepSeek loops for days: 237M tokens for just $3.48 — dogfoodarchitect · 2026-10-01
- Free course built from Cornell's GPU architecture workshop now shared publicly — idanbeck · 2026-10-01
- Qwen Flash Next MTP work resumes with official GGUF quants and llama.cpp PR — jacek2023 · 2026-10-01
- Undocumented Strata tip: set default sampling params via a sampling block in run config — KissMyShinyArse · 2026-10-01