FrontierCode benchmark grader called 'slop': penalizes good out-of-scope code changes
xeophon · x · 2026-09-23
The mystery of why models underperform on the FrontierCode benchmark has an answer: @rfxkairu points out FrontierCode penalizes out-of-scope code changes even when the changes are good. @xeophon confirms — "okay it's simpler, the grader is slop" — rejecting his earlier guess about oddly specified tasks forcing extra reasoning turns.
- Takeaway: how an agent coding benchmark's grader treats reasonable out-of-scope refactoring can directly distort model rankings
- A textbook counterexample for agent eval engineering: a poor grader makes benchmark results unreliable
More from coding & agent
- Mirage launches Tesseract, a video creative suite built for AI agents — _AustinCalvert_ · 2026-09-23
- Cloudflare launches Worker Previews: production-like environments per Git branch — dinasaur_404 · 2026-09-23
- qBotica: 3 founders to ~200 staff and 335% revenue growth by turning RPA into agentic AI software — rschmelzer · 2026-09-23
- A one-line prompt trick: 'Show me irresistible proof that you have fixed it' — jobergum · 2026-09-23
- Stripe's Harbor: AI-assisted prototyping tool produced 12,000 prototypes since May — dl_weekly · 2026-09-23
- Zvi on Opus 5.5 System Card: Failure Modes Are 'Largely Avoidable' via Prompting and Harness — TheZvi · 2026-09-23