Strata runs Qwen 3.8 Flash Next (125B) on a single RTX 4090 at 100 tok/s
snehesht · hn · 2026-10-04
A trending HN post showcases Strata, an open-source project claiming to run the 125B-parameter Qwen 3.8 Flash Next model on consumer hardware like the RTX 4090 at roughly 100 tokens/s. Code is available on GitHub for local deployment.
More from Infra
- Dev shares concurrency sweep method: TTFT, ITL and tok/s on 4xB200 for local models — TheZachMueller · 2026-10-04
- NVIDIA AIPerf docs go live: a package for performance-testing AI models — TheZachMueller · 2026-10-04
- One prompt freed 39.7 GB: Claude Code + ccmd MCP safely cleans dev caches — julsimon · 2026-10-04
- Community pushes llm-inference-bench as standard for local LLM inference speed measurement — TheZachMueller · 2026-10-04
- From one 3090 to 20 DGX Sparks: a home local-LLM cluster epic, 2.8T Kimi K3 at 20 t/s — ciprianveg · 2026-10-04
- Rural Queensland faces a $31B Anthropic datacentre, and locals aren't happy — nordicinst · 2026-10-04