Harness-IF Benchmark Finds Instruction Following Overestimated Across 12 Frontier Models

A new arXiv paper introduces Harness-IF, a benchmark for evaluating instruction following in coding agents, finding that all 12 frontier models are systematically overestimated by aggregate scores, since much apparent rule compliance is coincidental rather than genuine.

2026-08-14 ~ 2026-08-15 · 2 related posts