Dev reimplements DeepSeek v4.1 Flash, runs 1M context at 1-10 tok/s on a single RTX 4090
_xjdr · x · 2026-09-18
Developer xjdr implemented DeepSeek v4.1 Flash from scratch to understand the architecture, getting it to run at 1-10 tok/s (bs=1, ctx=1M+) at full precision (bf16/fp8/mxfp4 mix) on a single RTX 4090 with 64GB RAM and 1TB NVMe, depending on expert cache hits. His verdict: the architecture 'slaps' but is optimal for Grace Blackwell-class deployments, and he's impressed by DeepSeek's engineering under hardware constraints.
More from Infra
- Google Open-Sources Agent Substrate on GKE: 10x Density, 1,000+ Dormant Agents per Host — blaizedsouza · 2026-09-18
- Redditor crams six V100 GPUs into a standard full-tower case for local LLM inference — Odd_Caterpillar_2994 · 2026-09-18
- Crusoe raises $3.9B at $30.9B valuation to build data centers and modular AI factories — TechCrunch AI · 2026-09-18
- A 2.5-hour first-principles primer on the semiconductor supply chain worth your time — blaizedsouza · 2026-09-18
- YC F26's Dreamscale Labs Moves Robot AI Inference to the Cloud — ycombinator · 2026-09-18
- Community fine-tunes an MTP head for Bonsai 2 27B, ~1.25x inference speedup — cephaloform · 2026-09-18