CoRT: Optimizing LLM Fine-Grained Credit Assignment via Counterfactual Replay
_akhaliq · x · 2026-07-31
The paper introduces CoRT (Counterfactual Replay) to address the mismatch in rubric-guided LLM RL where rich criterion-level feedback is collapsed into a single response-level advantage and broadcast uniformly to all tokens.
By replaying responses with and without rubric criteria, CoRT turns policy-internal log-probability changes into normalized token-level credit weights without training an auxiliary scorer, changing the reward, or requiring extra generation/verifier calls. Across multiple models and instruction-following benchmarks, CoRT improves matched response-level GRPO by 4.4 points on average while remaining competitive with learned token-relevance methods.
Related event: ByteDance Proposes CoRT to Optimize LLM RL Credit Assignment(2 posts)→
More from Research
- Frontier AI Agents Fail at Open-Ended Research Despite 6 Days of Compute — sayashk · 2026-07-31
- AI Autonomously Tunes Quantum Devices Across Physical Platforms with 75 Diagrams — bravo_abad · 2026-07-31
- ICLR New Policy: Reviewers Must Disclose AI Tool Usage and Report Original Assessments — TuhinChakr · 2026-07-31
- Arena Introduces AutoEval: Minute-Level Ratings via Reward Models — arena · 2026-07-31
- LLM Analyzes 23,000 Western Books to Reveal 2,000 Years of Value Shifts — nwilliams030 · 2026-07-31
- MindForge: Significantly Boosting Small Models' From-Scratch Coding Abilities — centre-for-swe · 2026-07-31