456GB DeepSeek v4.1 runs locally at 40 tok/s with Threadripper + dual RTX 6000 hybrid setup

HankYeomans · x · 2026-09-22

A developer got DeepSeek v4.1 MXFP4 (about 456GB) running locally at 40 tok/s using a hybrid setup: 2x RTX 6000 Pro Max-Q for hot experts, a Threadripper Pro with 256GB CL32/6400MT RAM handling spillover experts, and Samsung 9100 NVMe for Engrams.

The goal was pushing non-GPU hardware to its limits and showing what big-memory workstation platforms can do for local LLM inference.

Original post →

More from Infra

Infra channel →