Microsoft collaboration TRACE tackles credit assignment for long-horizon agents
SharonYixuanLi · x · 2026-07-21
### Core idea Outcome-based RL works reasonably well when the task is short and the final answer is easy to verify, but it breaks down for long-horizon agents that may take dozens or hundreds of tool actions before producing an output. ### What TRACE does TRACE (**Turn-level Reward Assignment via Credit Estimation**) aims to assign rewards at the turn level **without**: - process labels - trained critics - Monte Carlo continuations - strong LLM judges The paper argues that credit assignment, not just scaling, is the real bottleneck for long-horizon agentic tasks. ### Why it matters The work suggests that if agents are going to handle long workflows reliably, reward design will likely need to move beyond a single final outcome signal and toward more direct per-turn supervision.
Related event: TRACE: A New Credit Assignment Method for Long-Horizon Agents(3 posts)→
More from coding & agent
- Kimi staff member builds a VR companion with Kimi Code K3 demo — dejavucoder · 2026-07-21
- OCR repo adds a JSON model directory to help agents pick the right model — strickvl · 2026-07-21
- Cross-agent system logs are dominated by questions and code proposals — nptacek · 2026-07-21
- Coding agents need better rules for when to read search summaries or full pages — RhubarbLarge2747 · 2026-07-21
- Notch says he may try vibe coding after struggling to hire good programmers — max_paperclips · 2026-07-21
- Seedance 2.0 keeps character consistency across 15+ shots with just 3 prompts — techhalla · 2026-07-21