CoRT: Optimizing LLM Fine-Grained Credit Assignment via Counterfactual Replay

_akhaliq · x · 2026-07-31

The paper introduces CoRT (Counterfactual Replay) to address the mismatch in rubric-guided LLM RL where rich criterion-level feedback is collapsed into a single response-level advantage and broadcast uniformly to all tokens.

By replaying responses with and without rubric criteria, CoRT turns policy-internal log-probability changes into normalized token-level credit weights without training an auxiliary scorer, changing the reward, or requiring extra generation/verifier calls. Across multiple models and instruction-following benchmarks, CoRT improves matched response-level GRPO by 4.4 points on average while remaining competitive with learned token-relevance methods.

Related event: ByteDance Proposes CoRT to Optimize LLM RL Credit Assignment(2 posts)→

Original post →

More from Research

Research channel →