Self-hosted LLM with SLERP-merged GRPO experts outperforms larger baseline, serving half of production traffic
t-tech · hf · 2026-09-02
t-tech shares how they trained separate GRPO experts on their real production request mix, then merged them via SLERP into a smaller self-hosted LLM.
The merged model outperforms a much larger baseline on instruction following, function calling, and internal tasks, while serving half of platform traffic at lower cost—demonstrating a practical path of post-training to your own traffic distribution.
More from Infra
- antirez: DeepSeek v4 Flash is still the king of local inference, now with vision — antirez · 2026-09-02
- Cheapest hardware to run Qwen3.8 27B at Q8? Ascend 310 out of stock, MI50 questioned — Snoo-2768 · 2026-09-02
- AI Infrastructure Night event in San Francisco — glcst · 2026-09-02
- Google signs 396 MW geothermal deal to power AI amid energy crunch — VraserX · 2026-09-02
- Local Model Suitability MCP: Cuts Costs via Local Inference — modelcontextprotocol · 2026-09-02
- OpenAI Engineer on Compilers 2.0: AI as Stochastic Optimizer — MikePFrank · 2026-09-02