Where does GPU spend actually go: training, inference, or idle reserved capacity?
Borges_Engineer · reddit · 2026-10-02
A Reddit discussion challenging the default assumption that ML compute costs mean training: inference runs all day, and reserved GPUs often sit idle between jobs because nobody wants to give up capacity. The author polls for rough splits (e.g., 30/50/20 training/inference/idle) and asks what actually worked to cut idle time — autoscaling, spot instances, shared queues, or social pressure. They also started a Discord for infra, serving, monitoring, and production-incident war stories.
More from Infra
- PyTorch's TLX-based JFA kernel beats FlashAttention-4 by 13% fwd, 50% bwd on B200 — PyTorch · 2026-10-02
- Fireworks shows numerical mismatch can collapse RL training in 25 steps on GLM and MoE models — sophiamyang · 2026-10-02
- $13B Baseten bets on open models, launches Base Labs research lab — baseten · 2026-10-02
- Kaigen details Unity benchmark setup: Runtime Speed with native C/C++ multithreading — gdechichi · 2026-10-02
- a16z: every $100 into AI buildout sends $50 to chips, $20 to power; 100+ charts — demian_ai · 2026-10-02
- Workload-aware inference: why batch LLM pipelines should plan queries like databases do — sh_reya · 2026-10-02