CER predicts terminal rewards early, cutting coding agent training tokens by 52.7%
shangbinfeng · x · 2026-10-08
The paper introduces Contextual Early Reward (CER): predicting the final reward from a trajectory prefix instead of waiting for full rollouts of long-horizon coding agents. CER adaptively synthesizes task- and stage-specific rubrics from summarized experience with related tasks, and uses the predicted reward for both test-time scaling and RL. On SWE-bench Verified, CER improves RM@8 over the strongest baseline by 4.2pp on Nemotron 3 Ultra and 2.0pp on Qwen 3.6 27B, matching baseline quality with only 15.3% of tokens; in RL it beats full-rollout TMax by 1.9pp while using 52.7% fewer online policy-and-judge tokens.
More from coding & agent
- DuckDB shows a database with no data: a few hundred KB catalog over your data lake — RexDouglass · 2026-10-08
- Dev open-sources smoke effect from Replit-built Drift game with agent prompt workflow — techartist_ · 2026-10-08
- Engineer Pushes Back on AI Code Analogy: Topology-Optimized Parts Look Cool but Are Unusable — rms80 · 2026-10-08
- Massive powers AI trade agents with x402 micropayments at a cent per request — kleffew94 · 2026-10-08
- A 4-month method for using Codex App as an information-discovery engine — jxnlco · 2026-10-08
- Handwritten text, AI-built interactions: a new golden age of technical explainers — srush_nlp · 2026-10-08