Model evals are getting harder: should we reward the smallest fix or the cleanest refactor?
hichaelmart · x · 2026-07-25
When a model fixes a bug, is fewer changes actually better?
The post argues that model evaluation is drifting into a messy question: should we prefer the programmer who fixes an issue with the fewest edits, or the one who refactors the surrounding code? The author says even humans can’t be judged with a simple binary rubric here, so model evaluators have an especially hard job.
The cited example suggests that Claude Opus 5 behaved differently under different effort settings: at low effort it made a single-character change, while at high effort it refactored the surrounding code, even though both runs ended up making the same effective fix. The implication is that output quality is not just about correctness, but also about style, scope, and maintainability.
More from Models
- A meme loops OpenAI, Anthropic, DeepSeek and Qwen through the same launch line — haider1 · 2026-07-25
- A thread argues that paperclip-maximizer fears came from a pre-LLM era of AI — repligate · 2026-07-25
- Anthropic ECI chart puts Claude Opus 5 at about 163.5 — scaling01 · 2026-07-25
- Grok 4.5 cache-read price cut lowers task cost by 25% — Baconbrix · 2026-07-25
- Reddit claims Claude Opus 5 beat Fable 5 in a 3D destruction test — Successful-Earth678 · 2026-07-25
- Claude may have stopped showing summarized thinking traces, users say — emollick · 2026-07-25