Only 15 Fused Kernels in Prefill: V4.1's Inference Stack and Minute-Scale SWA Cache
nrehiew_ · x · 2026-09-11
nrehiew sums up inference insights: heavily fused kernels (15 during prefill, 11 during decode), SWA-based KV cache reduction where SWA is cached only at message ends with a minute-scale TTL on DRAM, and a replay scheme bounding max replay by parameter S when caches are evicted. On post-training, he notes the team's philosophy is now completely data-focused rather than training-algorithm research.
More from Infra
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- Pentagon in talks to lend roughly $5 billion to AI cloud startup Fluidstack — vitaliychiley · 2026-09-11
- Eric Schmidt: AI may hit a money wall before a power wall — $1T capital needed — rohanpaul_ai · 2026-09-11
- SpaceX signs another AI compute deal: $1.11B per month, on track for $100B ARR — NinaDSchick · 2026-09-11
- Carmack: Jetson Thor's 128GB at 273GB/s is over-provisioned for real-time robotics — ID_AA_Carmack · 2026-09-11
- YC Demo Day startup touts ultra-pure diamond wafers for data centers, $160M in LOIs — ycombinator · 2026-09-11