DeepSeek V4 Flash Crashes During Prompt Processing on Dual Strix Halo RDMA Setup

WallabyFirm1159 · reddit · 2026-08-03

A developer attempted to run DeepSeek V4 Flash across two Strix Halo systems connected via RDMA using Mellanox ConnectX-3 cards and llama.cpp (ROCm).

While other models function properly on this exact setup, DeepSeek V4 Flash consistently crashes and dumps memory when prompt processing hits 4096 tokens. The author has tried debugging with Claude Code without success and is seeking community insights for this multi-node deployment issue.

Original post →

More from Infra

Infra channel →