Inference Engineering Is Just a Recipe: vLLM/SGLang, Replicas, Cache-Aware Routing
GabGarrett · x · 2026-09-03
Responding to hype about "learning inference engineering" (job postings up 163% YoY, $250K average salary), minhash argues there's little to learn: get GPUs, stand up vLLM or sglang (recipes included), run multiple replicas per node, add a cache-aware router (sglang's built-in one works), use LMCache so replicas see cache hits under load, and put a gateway in front of replicas on heterogeneous compute. The stack is rapidly becoming standardized.
More from coding & agent
- Fable 5.1 makes three.js sites: faster and sharper, but taste still matters — repligate · 2026-09-03
- SentinelX: open-source MCP agent lets LLMs control your Linux server via whitelisted commands — CarolusX74 · 2026-09-03
- Day 30 of U-BOT: fable 5.1 agent wrote a WiFi/BLE joystick controller — _Stocko_ · 2026-09-03
- First impressions of AmpCode: ChatGPT sub support, cloud-per-thread Orbs, and a hated mascot — carlosdponx · 2026-09-03
- One agent per user in a microVM: lessons from running a proactive AI friend in production — maritime_sh · 2026-09-03
- Writer has AI turn 10 handwritten pages into slides — and concludes his own prose reads like slop — cnakazawa · 2026-09-03