RTX 6000 Blackwell eGPU Crashes Under Heavy LLM Workloads, Requires Cold Reset
No-Paper-557 · reddit · 2026-08-11
A developer reported severe stability issues when running heavy LLM workloads via vLLM and llama.cpp on an RTX PRO 6000 Blackwell (96GB) housed in a Razer Core X V2 eGPU enclosure under Linux.
- Symptoms: While light workloads run fine, pushing the card with large contexts (e.g., 128K) or high VRAM utilization causes catastrophic crashes. The GPU logs Xid 175 (GSP RPC timeout) and Xid 154 (GPU reset required) errors, becoming entirely unresponsive until a hard cold power cycle.
- Potential Causes: The root cause remains unclear, with possibilities including Blackwell GSP driver bugs, USB4/Thunderbolt power management conflicts, or eGPU enclosure limitations.
- Current Mitigations: The developer has temporarily enabled PCIe runtime power control, turned on persistence mode, and downgraded the GPU power limit from 300W to 250W while seeking further community input.
More from Infra
- NVIDIA Nemotron 3.5 Lightning Hits DeepInfra with 1M Token Context — gharik · 2026-08-11
- China's DRAM Leader CXMT Joins MSCI China Index, Set to Lure Massive Fund Inflows — pstAsiatech · 2026-08-11
- OpenAI Pledges to Support New Power Generation and Grid Infrastructure in Texas — pstAsiatech · 2026-08-11
- MiniMax H3 Video Generation Benchmark: RTX 5090 Takes Under 5 Minutes — gabxav · 2026-08-11
- Open-Source Semantic LLM Cache PromptCache Cuts Costs by 80% with Sub-Millisecond Latency — tom_doerr · 2026-08-11
- Nvidia's HBM Reduction Is an Emergency Measure, Not a Compute Breakthrough — JOBhakdi · 2026-08-11