NVIDIA Paper Ranks Base Models by the One Edit That Fixes the Task

dair_ai · x · 2026-10-09

NVIDIA published a paper on choosing base models for coding agents, noting that base models are extremely hard to evaluate on agentic coding: across six base models run on SWE-bench Verified, five solved zero tasks.

Instead of end-to-end evaluation, the method isolates the single decisive step in a coding run. It takes tasks a strong post-trained agent already solved, replays its code edits one by one, and runs tests after each edit; the first edit that makes tests pass is the decisive edit. The base model is then given everything before that edit and must produce the fix.

Scoring covers three angles, including how likely the base model is to write the fix and whether it can pick the correct one. The authors frame this as a clever way to rank base checkpoints by their likely performance as coding agents after post-training.

Related event: NVIDIA Paper: Predicting Coding-Agent Performance From Base Models(2 posts)→

Original post →

More from coding & agent

coding & agent channel →