Is Self-Hosting LLMs on Preemptible Cloud Instances Cost-Effective? Reddit Discusses
Different-Monk5916 · reddit · 2026-08-15
A Reddit user explores using preemptible cloud instances (e.g., GCP L4, A100) to self-host LLMs with vLLM/LiteLLM, aiming to reduce inference costs. They plan to batch requests and tolerate longer completion times. Key concerns include instance reclaim handling, model loading overhead, storage costs, and overall feasibility for 20 hours/week usage. The post seeks community experience and advice.
More from Infra
- Sanders Threatens AI Pause, Nvidia Partners with Wall St on Compute-Backed Securities — PeterDiamandis · 2026-08-15
- Fireship: How edge ML cameras built a warrantless tracking network — Fireship · 2026-08-15
- Qwen Official Guide: Tuning reasoning depth and extending context to 1M tokens — solyarisoftware · 2026-08-15
- DeepSeek V4 Pro Requires Linux/WSL to Function Properly — teortaxesTex · 2026-08-15
- Tiny-Qwen update: PyTorch-native build for Qwen 3.8 27B model — No-Compote-6794 · 2026-08-15
- DeepSeek-V4-Flash on 4×AMD V620: 300K Context, 21 tok/s Generation — Thin_Pollution8843 · 2026-08-15