Strata runs a 125B model at 70tps on old DDR3 PCs with a cheap GPU
udmrzn · x · 2026-10-03
Strata (5.7k GitHub stars) is an open-source inference engine that runs the 125B-parameter Qwen3.8-Flash-Next on consumer hardware with one-click install for Windows/Linux. A cheap DDR3 machine with 64-128GB RAM plus a 8GB+ GPU hits 70+ tokens/s — more CPU cores means more speed. It exposes OpenAI/Anthropic-compatible local APIs, supports 128K context and image input, and can plug into coding agents with nothing leaving your machine.
More from Infra
- Just 10% of people using AI heavily could strain inference infra as HBM costs skyrocket — XFreeze · 2026-10-03
- Critics Slam a "Neocloud": Wrapped RunPod Months Ago, No Colocated DCs — knowrohit07 · 2026-10-03
- Analyst Bullish on AAOI as 800G/1.6T Ramp; 3.2T Not Mainstream Yet — BenBajarin · 2026-10-03
- Google Dropped Its Tier 2 Spend Gate That Pushed a Dev to OpenRouter — vivekhaldar · 2026-10-03
- AMD: Software Tuning on MI355X Plus vLLM 0.30.1 Cuts MiniMax-M3 Token Cost by 31% — ryanshrout · 2026-10-03
- 4-GPU server topology: one Gen5 switch covers it, benchmark all-reduce before buying three — knowrohit07 · 2026-10-03