Microsoft collaboration TRACE tackles credit assignment for long-horizon agents

SharonYixuanLi · x · 2026-07-21

### Core idea Outcome-based RL works reasonably well when the task is short and the final answer is easy to verify, but it breaks down for long-horizon agents that may take dozens or hundreds of tool actions before producing an output. ### What TRACE does TRACE (**Turn-level Reward Assignment via Credit Estimation**) aims to assign rewards at the turn level **without**: - process labels - trained critics - Monte Carlo continuations - strong LLM judges The paper argues that credit assignment, not just scaling, is the real bottleneck for long-horizon agentic tasks. ### Why it matters The work suggests that if agents are going to handle long workflows reliably, reward design will likely need to move beyond a single final outcome signal and toward more direct per-turn supervision.

Related event: TRACE: A New Credit Assignment Method for Long-Horizon Agents(3 posts)→

Original post →

More from coding & agent

coding & agent channel →