Samsung Pools KV Cache Using CXL
lauriewired · x · 2026-07-12
Samsung has published a paper/proposal detailing the offloading of KV-cache to a CXL memory pool.
- The experiment utilized an early CXL 2.0 ecosystem, older switches, and a relatively basic setup.
- By applying simple interleaving of the KV Cache across multiple CXL modules, the GPU supply remained stable, performing almost "like real DRAM".
- While the author notes this isn't SOTA, the implementation is highly reproducible, demonstrating that CXL already offers practical value for scaling inference memory.
More from Infra
- Nebius says SlimSpec speeds speculative decoding 8–9% without shrinking the vocabulary — Arindam_1729 · 2026-07-21
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- Chamath says open-sourcing Grok would push AI margins from models to infra and apps — Dan_Jeffries1 · 2026-07-21
- EU AI competitiveness is under pressure as firms double down on chips, ethics, and talent — nordicinst · 2026-07-21
- AI bottlenecks are shifting to memory, optics, yield control and power — thedealdirector · 2026-07-21