DeepSeek V4.1 Flash runs locally on 128GB Strix Halo via SSD streaming or two-node TCP/RoCE
antirez · x · 2026-09-23
Developer dcapitella added support for DeepSeek V4.1 Flash in antirez's DwarfStar local inference tool, running it on AMD Strix Halo (128GB unified memory) and testing two deployment paths: SSD streaming on a single 128GB machine, or distributing the model across two nodes over TCP or RoCE networking. A concrete demonstration of running large models off the data center on workstation-class hardware.
More from Infra
- The AI race becomes an efficiency race: fewer parameters, cheaper inference wins — ingliguori · 2026-09-23
- India starts commercial chip packaging as Tata builds $13.5B Dholera foundry — shashib · 2026-09-23
- Anthropic In Early Talks To Lease 1GW Of Data Center Capacity, TPU-Filled Sites Rumored — mark_k · 2026-09-23
- Hugging Face TRL Gets Major Speedup: AsyncGRPO Training Steps Up to 3.5x Faster — ben_burtenshaw · 2026-09-23
- Qualcomm ships two mobile chips with Hexagon NPU built for on-device MoE AI agents — emmanuelvivier · 2026-09-23
- Go.AI raises $85M for on-prem AI infrastructure serving banks and healthcare — emmanuelvivier · 2026-09-23