Own or rent? A practical guide to open-weights LLMs vs frontier APIs
spilldahill · reddit · 2026-09-07
A practical guide to deciding between self-hosted open-weights models and frontier APIs. Key points: a LoRA fine-tune of a 7B-8B model costs a few hundred dollars, after which per-token cost drops far below API pricing — a crossover point exists on the volume axis. Beyond cost: self-hosting cuts latency (200ms vs 2s matters for agent loops), satisfies data residency (GDPR/HIPAA), and avoids vendor lock-in from rate limits and model sunsets. Self-host when volume is high, the task is narrow, or data can't leave your perimeter; rent when volume is spiky or you're still exploring. Author discloses they work at Overmind.
More from Infra
- Compute financing risk will fall to hedge funds and commodity traders, not private credit — AccBalanced · 2026-09-08
- Solo dev ships Jenny, an MIT-licensed local LLM desktop app after 1.5 years — TangySword · 2026-09-08
- Cacheon launches GLM-5.3 kernel arena, paying up to 33 TAO daily to beat sglang — JosephJacks_ · 2026-09-08
- Johns Hopkins report urges strategic power islanding to cut blackout risk — philvenables · 2026-09-08
- HBF math: matching H200 bandwidth needs ~4,900 concurrent NAND planes — lauriewired · 2026-09-08
- BMO: record US power demand isn't lifting gas use — renewables and batteries are filling the gap — aronchick · 2026-09-08