Troubleshooting Qwen3.8 Flash Next on Mixed NVIDIA/AMD Hardware
Designer_Elephant227 · reddit · 2026-08-31
A user is attempting to run Qwen3.8 Flash Next locally on a mixed GPU setup (RTX 5070 Ti 16GB + Radeon AI PRO R9700 32GB + 96GB RAM) but is facing issues. They previously ran the 27b model successfully (q8 quantization, 256k context, 40tok/s). Current challenges include an AMD Pro driver unload bug causing system freezes and Claude incorrectly estimating the MoE model's VRAM requirements. The author is seeking advice on specific quantization choices, token speeds, launch flags, and handling the mixed NVIDIA/AMD environment.
More from Infra
- Verifying physical GPU assets: The challenge of financing compute infrastructure — AccBalanced · 2026-08-31
- Poll: 75% of Americans Oppose Local Data Centers — AndyMasley · 2026-08-31
- Dev Predicts Memory Crunch to Last Until Late 2027 — felpix_ · 2026-08-31
- Developer Warns Against 512GB Laptops Due to Dev Storage Needs — felpix_ · 2026-08-31
- China's CXMT reportedly achieves major breakthrough in advanced AI memory — Polymarket · 2026-08-31
- China's CXMT makes its first HBM3E chips, closing the AI memory gap — The Decoder · 2026-08-31