Infra engineers compare notes on when self-hosting LLMs beats paying for APIs
One_Mention_5385 · reddit · 2026-10-02
An infra engineer on Reddit is crowdsourcing real-world numbers on when self-hosting LLMs actually beats paying API bills, noting every blog post just says "it depends." He's asking teams that made (or abandoned) the switch for: the API spend that triggered the move, true total costs including GPUs, idle time and engineering hours, the most painful operational parts (cold starts, autoscaling, OOMs, model quality, 2am pages), and why teams that went back to APIs did so. He pledges to compile replies into a cost/decision writeup for the community.
More from Infra
- Cloudflare unveils open-source decision models Clef and a new RL fine-tuning platform — petrusenko_max · 2026-10-02
- Local AI roundup: 27B reasoning in 5.9GB, phone-class 35B, and dozens more — vramkickedin · 2026-10-02
- Liquid AI's Decision Model D1 Hits OpenRouter: Typed Answers With Probabilities at $0.04/M Input, $0 Output — maximelabonne · 2026-10-02
- Teenager tapes out a chip, rebuilds GPU interconnects with optics and poaches Nvidia veterans — ai · 2026-10-02
- huggingface_hub v2.1.0 ships: 17x faster downloads, job retries, rerun and port exposure — huggingface · 2026-10-02
- Gemini 4 isn't even out yet — but Google's TPU advantage is being underappreciated — haider1 · 2026-10-02