CER predicts terminal rewards early, cutting coding agent training tokens by 52.7%

shangbinfeng · x · 2026-10-08

The paper introduces Contextual Early Reward (CER): predicting the final reward from a trajectory prefix instead of waiting for full rollouts of long-horizon coding agents. CER adaptively synthesizes task- and stage-specific rubrics from summarized experience with related tasks, and uses the predicted reward for both test-time scaling and RL. On SWE-bench Verified, CER improves RM@8 over the strongest baseline by 4.2pp on Nemotron 3 Ultra and 2.0pp on Qwen 3.6 27B, matching baseline quality with only 15.3% of tokens; in RL it beats full-rollout TMax by 1.9pp while using 52.7% fewer online policy-and-judge tokens.

Original post →

More from coding & agent

coding & agent channel →