Harness-IF Benchmark: AI Coding Agents Don't Truly Follow All Rules

omarsar0 · x · 2026-08-14

While current coding agents follow user-defined rules (like those in AGENTS.md), they often do so because they were going to act that way anyway. To separate genuine instruction-following from coincidental behavior, researchers introduced Harness-IF.

Evaluation Method:

It scores 256 rules one at a time from execution evidence, then re-runs every task with the rule withheld across nine probe builds to identify which rules actually oppose the model's defaults.

Key Findings:

Original post →

More from coding & agent

coding & agent channel →