Benchmarks show coding agents edit code they shouldn't in 35-65% of cases; prompt framing is the lever
RunAI_Coder · reddit · 2026-09-21
The author ran 15 headless runs of "make one button blue" on a four-file page full of bait; every run stayed within one selector, showing small mocks say little since shared class names telegraph the edge. The real evidence comes from a benchmark of 200 SWE-Bench Verified issues where the real fix was pre-applied: current models in their own vendors' harnesses still edited code 35-65% of the time when the correct answer was an empty patch, and 87.1% of one model's failures touched code unrelated to the issue. Prompt wording matters: "fix the codebase" lowers correct abstention (65.0→56.5, 60.5→36.5), while "reproduce, then abstain if nothing's wrong" lifts it to 80.5 and 88.5 — at the cost of over-abstaining when a wrong patch needs replacing. Harnesses share the gap: edit auto-accept grants the whole working directory, path rules miss scripts that open files directly, and selectors inside allowed files are invisible to every layer.
More from coding & agent
- Addy Osmani shows Claude Code can auto-evaluate whether a plugin improves answers — addyosmani · 2026-09-21
- Devs say avoiding cache misses could boost Claude Code/Codex effective usage limits 10-20% — chaseleantj · 2026-09-21
- Solo dev's open-source AI workflow platform hits 38 deployments, now faces maintenance questions — Feathered-Beast · 2026-09-21
- Where should agent action authorization live? AI support teams debate refund safety policies — witty_queen123 · 2026-09-21
- Two years from learning Claude Code to $35k freelance income and a failed first startup — PratikKadam_ · 2026-09-21
- ostris ai-toolkit ships Qwen-Image-2.1 LoRA training support, no community LoRAs yet — reeight · 2026-09-21