GPU Inference Explained: Memory Bandwidth, Not Compute, Caps Tokens Per Second
abhijithneil · x · 2026-09-04
abhijithneil explains GPU inference with a factory metaphor: a GPU is a floor with thousands of workstations, and generating one token at a time hands it one-item-wide jobs. If you load all weights into memory, running N billion weights at 2 bytes each means 2N GB moved per token — bandwidth divided by that is your throughput ceiling, so GPU utilisation can be identical for 1 vs 100 users.
More from Infra
- AMD's Threadripper Halo Station packs 96 cores and 576GB of HBM3E — ccerrato147 · 2026-09-04
- AeroJEPA fluid foundation model joins NVIDIA's PhysicsNeMo ecosystem — ricardovinuesa · 2026-09-04
- Building a €2-2.5k local AI rig for legal RAG and agentic coding: hardware picks debated — whatyathinkk · 2026-09-04
- Dual 3090 owners debate adding more cards: bigger local models vs parallel instances — Blues520 · 2026-09-04
- NousResearch brings one-click local model setup to Hermes Agent on NVIDIA systems — lifebypixels · 2026-09-04
- Regulated-industry dev seeks AI Gateway with Okta SSO and runtime policy enforcement — IrrepressibleInk · 2026-09-04