A benchmark can say an agent solved the task while a nearby phrasing breaks it

bibryam · x · 2026-07-26

The post argues that benchmark scores can be misleading because a single phrasing may make an agent look “solved” even when performance collapses on nearby variants.

The attached chart illustrates the point: one query can hit F1 = 1.00 while aggregate F1 stays much lower across ambiguity levels. The takeaway is that evaluators should test the eval itself before trusting the reported score.

Original post →

More from Research

Research channel →