Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0%

antgroup · hf · 2026-07-21

What the paper says

Multi-teacher on-policy distillation for agentic language models can improve tool-call recall, but the paper finds it also shifts behavior toward over-calling: the student starts invoking tools in cases that should be answered directly.

Results

The main takeaway is that multi-teacher OPD should monitor not just how much teacher signal there is, but where it acts during generation.

Original post →

More from coding & agent

coding & agent channel →