Budget Hybrid Local Box for Qwen3.8-Flash-Next 4-bit: 96GB RAM Plus Used RTX 3090
Confident-Truth3607 · reddit · 2026-10-06
After testing 20 open models, the author picked Qwen3.8-Flash-Next (medium reasoning) — the only one that didn't invent config options when run with a doc-lookup harness — for a budget local build: Ryzen 5 9600 (€200), 2×48GB DDR5-5600 (€1199–1549, two sticks to avoid four-stick penalties), and a used RTX 3090 (€1150–1500), with 79GB of model in RAM and 20.6GB on GPU. Open questions: will 6 Zen 5 cores bottleneck 40 CPU MoE layers, is 17GB headroom enough, and can the Arc Pro B60 (€772) work via Vulkan/SYCL.
More from Infra
- NVIDIA report: 89% of telecom operators say open models are key to their AI strategy — nordicinst · 2026-10-06
- Dev's custom vLLM patch runs DiffusionGemma on DGX Spark, edging out the commercial API — vllm_project · 2026-10-06
- How a Solo Dev Stopped Local 7B Models from Hallucinating Bank Balances: Code Calculates, Model Summarizes — Revibed69 · 2026-10-06
- Cloudflare launches cf, an agent-first CLI to query observability data via the API — dinasaur_404 · 2026-10-06
- Building a local LLM agent stack on a 128GB Mac Studio: Reddit thread weighs inference layer options — DrainBramage · 2026-10-06
- Polars 2.0 ships: streaming by default, SQL first-class, beats DuckDB on TPC-H — banteg · 2026-10-06