DeepSeek V4.1 Flash infra: Siglip-style vision encoder and shadow indexer workers
nrehiew_ · x · 2026-09-11
nrehiew breaks down the training infra of DeepSeek V4.1 Flash:
- Vision encoder: Siglip-style; key insight is that all-gather of one modality's features overlaps with the other modality's compute; context parallelism for ultra-long multi-image sequences
- CSA 2 requirements: shadow indexers only do computation and pull updated weights from a master replica (similar to inference workers in RL); pipeline payload extensions, apparently metadata across pipeline stages
- Optimizer sharding: sharded across Engram tables, and Engram sitting in early layers enables prefetching while the vision encoder runs
More from Infra
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- Pentagon in talks to lend roughly $5 billion to AI cloud startup Fluidstack — vitaliychiley · 2026-09-11
- Eric Schmidt: AI may hit a money wall before a power wall — $1T capital needed — rohanpaul_ai · 2026-09-11
- SpaceX signs another AI compute deal: $1.11B per month, on track for $100B ARR — NinaDSchick · 2026-09-11
- Carmack: Jetson Thor's 128GB at 273GB/s is over-provisioned for real-time robotics — ID_AA_Carmack · 2026-09-11
- YC Demo Day startup touts ultra-pure diamond wafers for data centers, $160M in LOIs — ycombinator · 2026-09-11