Dev reimplements DeepSeek v4.1 Flash, runs 1M context at 1-10 tok/s on a single RTX 4090

_xjdr · x · 2026-09-18

Developer xjdr implemented DeepSeek v4.1 Flash from scratch to understand the architecture, getting it to run at 1-10 tok/s (bs=1, ctx=1M+) at full precision (bf16/fp8/mxfp4 mix) on a single RTX 4090 with 64GB RAM and 1TB NVMe, depending on expert cache hits. His verdict: the architecture 'slaps' but is optimal for Grace Blackwell-class deployments, and he's impressed by DeepSeek's engineering under hardware constraints.

Original post →

More from Infra

Infra channel →