1080 Ti beats RTX 6000 by 2.4x on dense-model inference despite 4x less bandwidth
EAccelerate_42 · x · 2026-09-22
A counterintuitive inference benchmark: loading a dense model that takes 80% of VRAM (rest for KV cache), the RTX 6000 (1,792GB/s bandwidth) manages 18.7 sweeps/s at 77GB per token (23 tok/s), while the 2017-era 1080 Ti (484GB/s) hits 44 sweeps/s at 8.8GB per token (55 tok/s) — 2.4x faster despite a quarter of the bandwidth.
Key takeaway: when estimating how fast a GPU runs a model, don't just compare memory bandwidth — check sweeps per second too, since per-token sweep sizes vary widely across models.
More from Infra
- A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4 — willcb · 2026-09-22
- AMD's market cap went from $2B to $1T in the 11 years since Lisa Su became CEO — xiaosun86 · 2026-09-22
- Blackstone poured ~$100B into data centers this year, telling investors this is not the dot-com bubble — JOBhakdi · 2026-09-22
- Modular Releases Free LLM Inference Handbook With 20+ Interactive Visualizations — carrycooldude · 2026-09-22
- Naveen Rao at All-In Summit: AI is extremely inefficient and every computer since 1945 is built wrong — NaveenGRao · 2026-09-22
- GKE Pod snapshots cut AI inference cold starts by 89%, loading 70B models in 37s — rseroter · 2026-09-22