NVIDIA Paper Ranks Base Models by the One Edit That Fixes the Task
dair_ai · x · 2026-10-09
NVIDIA published a paper on choosing base models for coding agents, noting that base models are extremely hard to evaluate on agentic coding: across six base models run on SWE-bench Verified, five solved zero tasks.
Instead of end-to-end evaluation, the method isolates the single decisive step in a coding run. It takes tasks a strong post-trained agent already solved, replays its code edits one by one, and runs tests after each edit; the first edit that makes tests pass is the decisive edit. The base model is then given everything before that edit and must produce the fix.
Scoring covers three angles, including how likely the base model is to write the fix and whether it can pick the correct one. The authors frame this as a clever way to rank base checkpoints by their likely performance as coding agents after post-training.
Related event: NVIDIA Paper: Predicting Coding-Agent Performance From Base Models(2 posts)→
More from coding & agent
- Designer spends a full day building a custom design tool with Claude to replace limiting Figma — oykun · 2026-10-10
- Coinbase CEO Brian Armstrong Ships Keyboard Shortcuts to Production Using AI Dev Tools — kleffew94 · 2026-10-10
- Open-EmbeddingGemma: pure PyTorch reimplementation of Google's 740M multimodal embedding model — KyeGomezB · 2026-10-10
- LLM Wiki, an open-source RAG alternative, hits 20.4k GitHub stars — tom_doerr · 2026-10-10
- Comma launches as an open-source 24/7 personal AI agent with no session limits — HeyToha · 2026-10-10
- Every ships a month of work in one week with AI and asks: what if we could publish 100 apps a week? — danshipper · 2026-10-10