Developer maps the entire inference + fine-tuning provider landscape with tradeoffs

Present-Jelly9941 · reddit · 2026-09-09

A developer who spent weeks shopping for inference + fine-tuning published a full landscape map. Key thesis: grouping matters more than names — vendors within a group are substitutable, across groups they aren't. Groups cover token-API+FT (Fireworks, Together), FT-first (Predibase, OpenPipe), managed deploy (Baseten, Replicate, Modal), raw GPU (RunPod, Lambda, CoreWeave), speed specialists (Groq, Cerebras), routers (OpenRouter), and Akka's cost-per-task routing. Self-hosting floor: vLLM/SGLang + LoRAX, trained with Axolotl/Unsloth/TRL. Per-token vs cost-per-task is a bet on workload narrowness, not a benchmark decision.

Original post →

More from Venture

Venture channel →