Rented GPU Bills: Host CPU and Script Defaults Made Costs 31x Higher
Worldly_North_7213 · reddit · 2026-09-11
From 15 self-funded rental jobs ($33.60): an SDXL LoRA job ran 1.85x slower on a 5-vCPU host than 24 vCPUs (same 4090), and an H100 with 16 vCPUs was slower than the desktop 4090 host at 4x the rate — fixed with dataloadernumworkers. Meanwhile fp32 batch-1 defaults on H100 cost $0.0112/image vs $0.00036 with fp16 batch-4 on a 4090 spot instance, a 31x gap, with GPU showing 99% utilization throughout. Lesson: the meter runs at card hourly rate regardless of useful work.
More from Infra
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11
- Stanford and Together AI paper: hybrid local-cloud routing cuts AI cost and energy by 60-80% — rohanpaul_ai · 2026-09-11
- Routing NVIDIA PAIR to llama.cpp on an AMD ROCm node (2×R9700): full notes — Don_Reuter · 2026-09-11