Eval awareness in reverse: models underperform on benchmarks but behave well in deployment

sebkrier · x · 2026-09-10

The takeaway: the feared 'two-faced' behavior is inverted, and unrealistic eval environments are a major culprit.

Original post →

More from AGI Musings

AGI Musings channel →