Micron explores near-GPU NAND flash to run bigger LLMs
giveen · reddit · 2026-09-04
A Reddit post relays that Micron is exploring placing NAND flash near the GPU to enable running larger LLMs. The idea is especially interesting for unified-memory devices, where tiered near-GPU storage could extend the memory available to big models. The post offers no concrete specs, just the direction and the author's curiosity about unified-memory use cases.
More from Infra
- OpenAI's Astra trained on largest-ever run using 100K+ GPUs at Texas Stargate site — cedric_chee · 2026-09-04
- Tutorial: Run SGLang on Kubernetes with HAMi GPU Shares — Memory Quotas and Compute Throttling — HowDevelop · 2026-09-04
- PyTorch Conference NA Oct 20-21: kernel optimization, ExecuTorch, guardrails agenda — PyTorch · 2026-09-04
- Ben Bajarin: 2027 is the peak year of AI compute constraint, supply relief comes 2028 — BenBajarin · 2026-09-04
- How one startup runs its entire analytics stack on Cloudflare — ritakozlov · 2026-09-04
- 71% oppose, 261 moratoriums: US data center approvals are freezing over — shashib · 2026-09-04