Local Qwen3.8-flash Test: Strix Halo and a 94GB-Modded RTX 3090 Both Hit ~50 tok/s
drdanielbender · x · 2026-10-09
A side-by-side local inference test of qwen3.8-flash-next across two setups:
- AMD Strix Halo 128GB (using @high52weeks' Halogen)
- RTX 3090 desktop with 94 GB of memory (using @coldniko's Strata)
- Both setups reach roughly 50 tok/s and feel smooth
- The author uses them to power a Hermes agent for privacy-sensitive tasks, fully local
A direct hardware/software combo reference for developers eyeing local deployment and privacy-focused workflows.
More from Infra
- Broadcom CEO Hock Tan: Competing Outside Your Lead Market Gets You 'Your Ass Kicked' — zephyr_z9 · 2026-10-09
- Deno joins Cloudflare — a core engineer on following the team over — cnakazawa · 2026-10-09
- Oxide Computer raises $445M Series D as AI demand drives enterprises to rethink owning vs renting compute — Sethwinterroth · 2026-10-09
- Datology AI launches Curation Studio, pitching data quality as the ultimate compute multiplier — schwarzjn_ · 2026-10-09
- Impactful Scheduling for GPU Clusters: Inside AI2's New Scheduler — Hugging Face Blog · 2026-10-09
- The 'Montreal Premium': Same GPUs Cost 20%+ More Per Hour in Canada Than the US — RichardsonDx · 2026-10-09