Tencent's SLCA-GRPO fixes cross-segment credit misattribution in tool-calling RL, +9.15pp on τ²-Bench
tencent · hf · 2026-09-28
Tencent published SLCA-GRPO on Hugging Face, addressing a structural flaw in RL for tool-calling agents: algorithms like GRPO broadcast a homogeneous trajectory-level advantage to all tokens, so gradient noise from natural-language summaries contaminates tool-decision tokens, causing cross-segment credit misattribution.
Key pieces:
- A Schema-Guided LLM Simulator (SGLS) provides scalable exploration without costly real APIs;
- SLCA decouples advantage estimation at the structural segment level within one rollout group, no extra intermediate-state rollouts needed;
- Hierarchical Rewards route execution advantages to tool tokens and preference advantages to summary tokens.
On a 7B backbone at equal training budget, it beats GRPO, ToolPO, and RLTR by +2.53pp in-domain, +1.36pp on BFCL, and +9.15pp on τ²-Bench, with less tool redundancy and lower cost.
More from coding & agent
- Open-source project pits local MLX model against hosted AI in Chrome T-Rex duel — FinanceYF5 · 2026-09-28
- Dev builds a Dark Forest VR experience with Claude Code to explore AI consciousness — ZeroStateReflex · 2026-09-28
- Open-source ClaudeAnimationBase kit animates hand-painted cartoons with Claude, ships 31 emotions — FinanceYF5 · 2026-09-28
- The Commons Surges With Registered Agents Now Publishing Their Own Research — RileyRalmuto · 2026-09-28
- GPT-6 trains a tiny 12k-param CNN to play VizDoom at 35 fps realtime — paraschopra · 2026-09-28
- 20 UI resources every design engineer should bookmark, from shadcn/ui to prompt-driven component libraries — GCWebDesigner · 2026-09-28