Is Self-Hosting LLMs on Preemptible Cloud Instances Cost-Effective? Reddit Discusses

Different-Monk5916 · reddit · 2026-08-15

A Reddit user explores using preemptible cloud instances (e.g., GCP L4, A100) to self-host LLMs with vLLM/LiteLLM, aiming to reduce inference costs. They plan to batch requests and tolerate longer completion times. Key concerns include instance reclaim handling, model loading overhead, storage costs, and overall feasibility for 20 hours/week usage. The post seeks community experience and advice.

Original post →

More from Infra

Infra channel →