FrontierCode benchmark grader called 'slop': penalizes good out-of-scope code changes

xeophon · x · 2026-09-23

The mystery of why models underperform on the FrontierCode benchmark has an answer: @rfxkairu points out FrontierCode penalizes out-of-scope code changes even when the changes are good. @xeophon confirms — "okay it's simpler, the grader is slop" — rejecting his earlier guess about oddly specified tasks forcing extra reasoning turns.

Original post →

More from coding & agent

coding & agent channel →