Strata open-source engine runs 125B Qwen3.8-Flash-Next on gaming PCs, 5k stars in 8 days
alex_verem · x · 2026-10-02
Strata, an open-source inference engine built by a solo developer, runs the 125B-parameter Qwen3.8-Flash-Next on ordinary gaming PCs and hit 5,000+ GitHub stars in 8 days. Install is one double-click, or one prompt in Claude Code/Cursor. On an RTX 5070 it generates faster than you can read, supporting chat, coding, image input, and OpenAI/Anthropic-compatible local APIs for other apps and coding agents. Requirements: 12GB VRAM, 32GB RAM, 80GB disk. Free, no account, fully local.
Related event: Strata: Running a 125B-Parameter Model on a Gaming PC(6 posts)→
More from Infra
- mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking — StasBekman · 2026-10-03
- Measured on B200: nvfp4 is ~9% more efficient than mxfp4 with higher accuracy — pick nvfp4 on Blackwell — StasBekman · 2026-10-03
- LithosAI launches LithosBox: millisecond snapshot-and-fork sandboxes for AI agents — JiaZhihao · 2026-10-03
- RTX Spark laptops and mini desktops rumored Oct 7 launch, $1800-$2900 with 24GB-128GB — Porespellar · 2026-10-03
- Wish list: a Qwen4 27B with 100B+ Engram offloaded to RAM and NVMe for local users — casper_hansen_ · 2026-10-03
- State of Local AI 2026: one gaming GPU now matches the world's best model from Feb 2026 — Scobleizer · 2026-10-03