Analog chip runs LLM attention 100x faster than H100 using 70,000x less power, Nature paper claims
anselm · x · 2026-09-25
- Researchers published a Nature paper on GainCellAttention, an analog in-memory computing architecture claimed to run LLM attention 100x faster than an H100 while using 70,000x less power.
- The pitch: every AI chip today inherits the von Neumann bottleneck — compute and memory are physically separated, and shuttling data back and forth can cost up to 10,000x more energy than the math itself.
- Using emerging "gain cell" hardware, the architecture performs the attention computation directly inside memory, eliminating the data movement bottleneck at its root.
- Still a research-stage result, but if it matures it could reshape inference economics; watch for follow-up on scalability and precision.
More from Infra
- Lambda engineer shares local inference build rule: 27B models need 24-32GB VRAM — TheZachMueller · 2026-09-25
- Pokee AI demos 36B agent model running fully local on Snapdragon X2 Elite with 32GB RAM — Kyrannio · 2026-09-25
- AMD to present MXFP8 pretraining scaling on 1K+ MI355X GPUs at PyTorchCon 2026 — PyTorch · 2026-09-25
- AI energy startup Parallax launches with $117m from Founders Fund, Lux, Greylock and others — graceisford · 2026-09-25
- Nebius/WEKA benchmark: shared KV cache lifts agentic inference throughput 2.4x with 93% hit rate — AccBalanced · 2026-09-25
- Burkov's TP Weekly #179: GPU rent vs buy, llm-d serving 753B model at 5-10x lower cost — burkov · 2026-09-25