Post-training, custom spec decoding and vLLM tuning: a hands-on inference cost-saving playbook
dhruv2038 · x · 2026-09-02
A first-hand account of deploying custom models for a large evohq customer over one week:
- Analyzed traffic patterns, set up evals, and ran extensive ablation experiments across multiple models
- When savings were insufficient, moved to post-training: SFT first, then additional RL policies
- End-to-end inference optimization: custom speculative decoding models fitted to the customer's data distribution, vLLM config tuning, quantization
- Also handled compute allocation, inference capacity planning, and warm-up based on traffic profiles
- The whole pipeline was orchestrated by an in-house autoresearch / AI engineer
The team is rapidly onboarding inference providers for growing demand and is looking for customers spending $20k+/month on agentic/AI workloads.
More from coding & agent
- Giving AI agents their own inbox is architecturally wrong, Reddit thread argues — Creamy-And-Crowded · 2026-09-02
- Merge launches enterprise AI governance tool enforcing model routing and spend rules — shensi · 2026-09-02
- Weaviate Ask Mode Adds Configurable Evaluation to Trade Latency vs Verifiability — CShorten30 · 2026-09-02
- Building the Ultimate Agent Harness for Kimi K3: The Model Is No Longer the Bottleneck — VibeMarketer_ · 2026-09-02
- GLM 5.2 slug references spotted in Google Antigravity CLI, hinting at integration — gaganghotra_ · 2026-09-02
- Design lead ships 12 PRs in a week: AI is erasing the designer-engineer gap — talkaboutdesign · 2026-09-02