Devs call for standardized "model performance across harnesses" evals

zainhas · x · 2026-09-10

A developer argues that the same new model can perform wildly differently depending on which harness you use, making it impossible to tell whether you're getting a frontier model or something a year old — and calls for the industry to routinely publish model performance across harnesses so evaluation setups become transparent.

Original post →

More from Models

Models channel →