How to locally run a Qwen 3.8 Next model at Claude Opus level for under $4,000

ironicstatistic · reddit · 2026-09-17

A Reddit user asks for deployment advice on building a dual-GPU local server to run a Q3 quant of Qwen 3.8 Next (85GB with KV cache) at 256k context, targeting Claude Opus-class intelligence with 500 tok/s prefill and 20 tok/s decode on a $4,000 budget.

Four candidate builds:

His thesis: with optimizations like n-gram spec decoding, mid-tier Opus-class models are becoming the realistic local-deployment target.

Original post →

More from Infra

Infra channel →