NVIDIA and Solidigm Rewrite Storage Rules for AI's KV Cache

shashib · x · 2026-08-02

NVIDIA and Solidigm are co-designing solid-state drives that sit directly inside the GPU memory hierarchy to handle the growing KV cache in LLMs.

The core premise is that with high bandwidth memory (HBM) costing around $10,000 per terabyte, pushing overflow context to flash storage is cheaper. This new architecture allows the storage to drop data bytes because, in this specific scenario, a dropped byte simply triggers a recompute rather than a catastrophic data loss event.

Original post →

More from Infra

Infra channel →