NVIDIA-led paper predicts coding-agent post-training gains from base models
heghbalz · x · 2026-10-09
A new arXiv paper, "Before They Can Solve: Predicting Post-Training Coding-Agent Performance from Base Models," tackles how to tell which base checkpoint merits an expensive round of agentic post-training.
Key points:
- End-to-end pass@$K$ poorly fits agentic coding: many base checkpoints can't reliably produce well-formed tool calls, and short-horizon evals sidestep multi-step state coherence.
- The authors use successful post-trained agent trajectories as a lookahead signal: replaying each trajectory and rerunning tests after every code-changing step identifies the "decisive step" — the first step whose cumulative patch flips the repo from failing to passing.
- Built on a coverage principle for agentic traces, the method predicts post-training performance directly from base-model behavior.
The 22-author list includes Ashish Vaswani, Bryan Catanzaro, and other NVIDIA researchers. Sharif Samee (rosinality) highlighted it as an interesting research direction beyond simply adding more agentic trajectories.
More from coding & agent
- Comma: open-source 24/7 personal agent that runs background tasks to completion — FellMentKE · 2026-10-09
- qwen-code v0.25.1-preview.1 ships agent, CLI and CI fixes — qwen-code-review-bot · 2026-10-09
- ByteDance merges TRAE product lines; blogger calls it the closest thing to Codex — vista8 · 2026-10-09
- 20 GitHub repos turn Claude into a motion design studio — Roger_M_Taylor · 2026-10-09
- TinyJoin v0.6 now beats SQLite and PGlite on 20 benchmarks, stays smaller — Vjeux · 2026-10-09
- Riviera lets you fix how agents use your API without redeploying the backend — Rodrigocruuz · 2026-10-09