Tencent's SLCA-GRPO fixes cross-segment credit misattribution in tool-calling RL, +9.15pp on τ²-Bench

tencent · hf · 2026-09-28

Tencent published SLCA-GRPO on Hugging Face, addressing a structural flaw in RL for tool-calling agents: algorithms like GRPO broadcast a homogeneous trajectory-level advantage to all tokens, so gradient noise from natural-language summaries contaminates tool-decision tokens, causing cross-segment credit misattribution.

Key pieces:

On a 7B backbone at equal training budget, it beats GRPO, ToolPO, and RLTR by +2.53pp in-domain, +1.36pp on BFCL, and +9.15pp on τ²-Bench, with less tool redundancy and lower cost.

Original post →

More from coding & agent

coding & agent channel →