DeepSeek V4 Flash Crashes During Prompt Processing on Dual Strix Halo RDMA Setup
WallabyFirm1159 · reddit · 2026-08-03
A developer attempted to run DeepSeek V4 Flash across two Strix Halo systems connected via RDMA using Mellanox ConnectX-3 cards and llama.cpp (ROCm).
While other models function properly on this exact setup, DeepSeek V4 Flash consistently crashes and dumps memory when prompt processing hits 4096 tokens. The author has tried debugging with Claude Code without success and is seeking community insights for this multi-node deployment issue.
More from Infra
- Open-Source System Serves VLA Models to 10+ Robots on a Single GPU — danfei_xu · 2026-08-03
- MiniMax H3 Tested: 8s Anime Video Generation on AMD GPU — klemze · 2026-08-03
- From 3D Gaming to AI Dominance: How NVIDIA Seized the Future of Computing — TinfoilTricorn · 2026-08-03
- Defending AI's Thirst: Are Data Centers Really Draining More Water Than Agriculture? — joshwhiton · 2026-08-03
- AI Chip Startup OLIX Raises $312M Series B at $3.3B Valuation — matthewclifford · 2026-08-03
- Run Local LLMs on Mac Easily: llama-macos Offers One-Click Server and WebUI — mervenoyann · 2026-08-03