Troubleshooting Qwen3.8 Flash Next on Mixed NVIDIA/AMD Hardware

Designer_Elephant227 · reddit · 2026-08-31

A user is attempting to run Qwen3.8 Flash Next locally on a mixed GPU setup (RTX 5070 Ti 16GB + Radeon AI PRO R9700 32GB + 96GB RAM) but is facing issues. They previously ran the 27b model successfully (q8 quantization, 256k context, 40tok/s). Current challenges include an AMD Pro driver unload bug causing system freezes and Claude incorrectly estimating the MoE model's VRAM requirements. The author is seeking advice on specific quantization choices, token speeds, launch flags, and handling the mixed NVIDIA/AMD environment.

Original post →

More from Infra

Infra channel →