Where does GPU spend actually go: training, inference, or idle reserved capacity?

Borges_Engineer · reddit · 2026-10-02

A Reddit discussion challenging the default assumption that ML compute costs mean training: inference runs all day, and reserved GPUs often sit idle between jobs because nobody wants to give up capacity. The author polls for rough splits (e.g., 30/50/20 training/inference/idle) and asks what actually worked to cut idle time — autoscaling, spot instances, shared queues, or social pressure. They also started a Discord for infra, serving, monitoring, and production-incident war stories.

Original post →

More from Infra

Infra channel →