HBF math: matching H200 bandwidth needs ~4,900 concurrent NAND planes
lauriewired · x · 2026-09-08
lauriewired breaks down High-Bandwidth Flash: matching an H200's 4.8TB/s bandwidth requires 4,900 NAND planes reading 4KB blocks concurrently, so the real question is whether parallelism lives in software or hardware. Huawei's FLINT uses a reactive SRAM-based burst controller (512 planes, reads only, good for weight loading), while the FlashAccel paper co-designs software/hardware with plane→megaplane→hyperpage (4,608 planes) coalescing usable as KV-cache. Verdict: raw HBF will be ugly, first useful versions FLINT-like, and FlashAccel-style designs the long-term winner.
Related event: Researcher Pours Cold Water on High-Bandwidth Flash Hype(2 posts)→
More from Infra
- NVIDIA's new Sol-H3 fast inference method for H3 awaits a ComfyUI port — krigeta1 · 2026-09-08
- Walking the AI rack optical stack: InP substrates and silicon photonics as the cleaner bet — demian_ai · 2026-09-08
- Running dual RX 7900 XTX on X570/X870 Taichi for local LLM inference: is x8/x8 enough? — espece-de-bon · 2026-09-08
- Hugging Face teases WebGPU inference engine with 5-10x speedups on Transformers.js — nicodotdev · 2026-09-08
- Meta to deep-dive recommendation inference systems at PyTorch Conference 2026 — PyTorch · 2026-09-08
- Memory crunch hits home: 4TB portable SSD prices stun as AI reprices the storage stack — demian_ai · 2026-09-08