LLM inference may already beat human inference on energy per token
jd_pressman · x · 2026-07-21
The post argues that LLM inference may already be more energy-efficient than human inference on a per-token basis, at least for short-context generation.
- A rough estimate uses Epoch’s inventory of about 48,000 NVIDIA NVL72 racks, roughly 5.8 GW of rack power, and divides by output capacity to get about:
- 0.3 J/token for short context
- 1 J/token for medium context
- 12 J/token for long context
- After accounting for cooling and power overhead, the author says the figures rise somewhat, but the core comparison still stands for short outputs.
- Using the tweet’s 12 W brain figure and 3–4 tokens/s gives about 3–4 J/token; using the more conventional 20 W whole-brain budget gives 5–7 J/token.
- Their conclusion: modern batched inference is already more energy-efficient per short-context token than the proposed human baseline, while very long-context inference is still somewhat worse.
- They also note an important caveat: Kimi K2.6 has 1T stored parameters, but only 32B active per token, not the roughly 360B active parameters assumed in the quoted comparison.
Related event: Epoch AI: LLM Query Energy Consumption Drops Below Human Brain's(2 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11