Running a 100B+ Qwen3 model locally on 64GB RAM: good vibes, short 192K context
lxfater · x · 2026-10-04
The author deployed a quantized 100B+ parameter Qwen3 model locally on a 64GB RAM machine. First impressions are positive — the new UI looks noticeably better — with the only drawback being the 192K context window. They're now looking for real use cases for their new local powerhouse.
More from Infra
- Local LLM benchmarks are mostly noise: c=1 tokens/s hides real concurrency performance — TheZachMueller · 2026-10-04
- Ahmad Osman on MTS Live: how local AI is becoming competitive, DGX Station in tow — basedjensen · 2026-10-04
- Building a Miro clone 5x on 3 local rigs: tokens/sec is useless, thinking variance hits 5x — julianharris · 2026-10-04
- RAM prices nearly quadruple in two years as users chase 256k contexts — lxfater · 2026-10-04
- TPU cost per million tokens beats NVIDIA Blackwell, giving Google an edge — cgarciae88 · 2026-10-04
- Strata calibrate nearly tripled decode speed: 256K context on a 16GB GPU — MoonsvnLyn · 2026-10-04