Self-hosted LLM with SLERP-merged GRPO experts outperforms larger baseline, serving half of production traffic

t-tech · hf · 2026-09-02

t-tech shares how they trained separate GRPO experts on their real production request mix, then merged them via SLERP into a smaller self-hosted LLM.

The merged model outperforms a much larger baseline on instruction following, function calling, and internal tasks, while serving half of platform traffic at lower cost—demonstrating a practical path of post-training to your own traffic distribution.

Original post →

More from Infra

Infra channel →