Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models
Tan Yu, Alexander Bukharin, Khushi Bhardwaj, Jennifer Williams, Zirui Liu, Jonathan Lingjie Li, Soumye Singhal, Joseph Jennings, Sanjeev Satheesh, Yash Jain, Ashish Vaswani, Venkat Krishna Srinivasan, Matthew Papakipos, Hyunwoo Kim, Jian Zhang, Oleksii Kuchaiev, Markus Kliegl, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jonathan Cohen, Jiantao Jiao
cs.AI, cs.SE
2026-10-08
Probing base checkpoints at the decisive step of agent trajectories tracks post-trained SWE-bench Verified rankings at ρ up to 0.964, vs 0.830 for the best bounded benchmark.
The most expensive decision in building a coding agent comes first: which base checkpoint goes into post-training. If SFT and RL reveal the wrong pick, the compute is already spent. Existing evaluation fails at this decision point from both directions.
End-to-end benchmarks such as SWE-bench Verified drop a base model into a tool-use harness, and the floor swallows it: five of six base checkpoints the paper tests score 0.0 pass@1, and the single non-zero score misleads, since that model's post-trained family ranks only fourth of six. Base models cannot reliably emit well-formed tool calls. Bounded single-shot benchmarks do run, but their predictive power is scattered: across ten public base/post-trained pairs, Spearman ρ with post-trained SWE-bench Verified pass@1 runs from 0.830 on RepoBench XFirst to -0.394 on HumanEval. Range compression is part of it; MBPP spreads only 67.2-86.0 across these ten models while the downstream target spans 38.8-80.6.
The starting point is the coverage principle: recent work finds that RLVR mostly sharpens solutions a base model can already sample rather than creating new ones. If post-training selects from the base distribution, the quantity to measure is the probability mass a base model keeps on successful agent behavior.
Where to measure is decided by the task's own verifier. Successful trajectories from frontier post-trained models (GPT-5.6-Sol, Opus5, Kimi-K3, running mini-swe-agent) are collected on DeepSWE, SWE-Pro, and SWE-bench Verified; replaying each trajectory, at every code-changing step the cumulative patch is extracted and the task's tests are run. The first step that flips the tests from failing to passing is the decisive step, and its action is the golden action, certified by the verifier rather than hand-written like the benchmark's gold patch. Earlier steps do not resolve the task and later ones may be unrelated churn, so only this step is worth probing.
Three probes run at that step, none requiring the base checkpoint to drive the harness:
Verifier executions are the scarce resource, so the decisive step is located by bisection in logarithmic calls per trajectory. The first two probes are static, one forward pass per item.
Ten public base checkpoints (three Nemotron-3 sizes, Qwen-3.5-35B, two DeepSeek-V4, Kimi-K2, GLM-4.5-Air, Tencent Hy3, Gemma-4-26B), scored against their post-trained families' SWE-bench Verified pass@1 (38.8 to 80.6):
| Probe (built on DeepSWE trajectories) | SWE-bench Verified ρ | Multilingual ρ | Terminal-Bench 2.1 ρ |
| Decisive-Action BPB | 0.964 | 0.976 | 0.833 |
| Patch MCQ | 0.867 | 0.927 | 0.796 |
| Prefix pass@K (K=32) | 0.951 | 0.945 | 0.930 |
The best bounded baseline reaches 0.830 (RepoBench XFirst; CRUXEval-O 0.758, MBPP 0.091, HumanEval -0.394). The result does not depend on trajectory source: probes built from SWE-bench Verified's own trajectories still reach 0.915, 0.903, and 0.906, and the cross-domain DeepSWE source actually agrees more closely, so same-benchmark leakage is not carrying the result.
Ablations back each design choice. Averaging BPB over all steps drops ρ from 0.964 to 0.915 and reverses two checkpoint pairs (DeepSeek V4 Flash vs Nemotron Ultra, Kimi K2 vs Qwen 3.5 35B). Scoring Patch MCQ options independently instead of jointly drops ρ from 0.867 to 0.665. A budget of K=16 captures nearly all the agreement (0.952, against 0.951 at K=32).
In a controlled SFT experiment (three base checkpoints, one recipe, 1,000 steps, 16.8B tokens), post-SFT scores of 28.5, 53.1, and 61.2 follow the ordering all three probes produce.
For checkpoint selection, the static probes need one forward pass per item and never touch the environment, yet rank checkpoints in close agreement with post-trained performance. That is a directly usable filter. The framework is the bigger prize: any agentic coding benchmark with successful trajectories and an executable verifier can be turned into a base-model evaluation, including future ones. Post-training potential becomes a signal you can monitor during training rather than a verdict you wait for, the same move as selecting pretraining data with loss correlations.
Honest framing: this is a correlation result on ten points, a ranking tool rather than a calibrated predictor. It says which checkpoint deserves the investment, not what score the investment will reach.
The paper's own list: ten checkpoints, no confidence intervals, descriptive point estimates; downstream scores come from different labs' post-training recipes and harnesses; the controlled SFT experiment covers only three checkpoints, and agreement weakens for some probes under other trajectory sources.
Transfer is uneven. On Terminal-Bench 2.1 the two static probes fall to 0.833 and 0.796; Kimi K2 and Gemma 4 already rank differently across the two benchmarks and the probes side with the SWE view. Covering command-line tasks likely needs command-line trajectories, which the paper leaves as future work.
Reading closely raises more concerns. Golden actions and distractors are sampled from specific frontier models: decisive-step certification settles what works, but the prefix itself carries those models' voice, so part of what the static probes measure may be likelihood of a particular style rather than problem-solving ability. Prefixes exceeding the context limit are dropped, and the longest prefixes correspond to the most-explored instances, a selection effect. Patch MCQ's absolute accuracy is modest (ten-model means of 0.411 to 0.504 against 0.250 chance): enough for ranking, weak per-item interpretability.