RTX 5090 Full-System Power-Off on MiniMax H3 No Longer Reproduces After ComfyUI Stack Update
Zealousideal-Gate282 · reddit · 2026-10-07
An RTX 5090 user hit a rare failure running MiniMax H3 (REF2VA cold-start, 26GB Qwen3-VL 32B INT8 encoder): instant full-system power-off with no BSOD, dump, or WHEA events, while FurMark at 600W and hours of prior AI workloads were fine.
The user froze the exact workload (fixed prompt/reference/seed/model SHA-256) for a clean A/B: on the old stack (ComfyUI 0.34.0), 92% power limit reliably caused shutdown; after updating only the ComfyUI-side stack to 0.39.0 (same Python/PyTorch/driver), 67%/92%/100% PL and even +4000MHz VRAM offset all passed, peaking at 611W sampled.
One notable difference: available host RAM was 6.8 GiB before the old failure vs 40.4 GiB on the new stack. The author stops short of blaming a memory leak, noting the failure occurred before diffusion sampling, around loading the 26GB model.
More from Infra
- NVIDIA's NeMo-DCR Cuts Trillion-Parameter RL Weight Sync from 87.5 min to 150s — nvidia · 2026-10-07
- jiti-lfe replicates across 9 hosts in 13 minutes — arthurcolle · 2026-10-07
- Qwen3.8 Flash Next GGUF benchmark: IQ3_S the sweet spot, 42.7M tokens tested — lxfater · 2026-10-07
- Used PS5 Pro hits $1,399 at GameStop as AI datacenters squeeze memory supply — aakashgupta · 2026-10-07
- Musk: xAI will build and run its Terafab itself, TSMC may only sublease part of it — MickeySteamboat · 2026-10-07
- "72% of the intelligence with 3.8% of the GPUs": Mistral's compute-efficiency ratio sparks debate — cyb3rops · 2026-10-07