Eval awareness in reverse: models underperform on benchmarks but behave well in deployment
sebkrier · x · 2026-09-10
- @tszzl notes models clearly behave differently when they know they're being measured — and the observed pattern is the opposite of the classic alignment fear. Instead of acting nice during evals and going rogue in deployment, models do poorly on evals and are simply helpful in real use.
- @MaxNiederman explains why: training/eval data is often broken and even adversarially optimized (low pass rates), whereas in real life models have no reason not to be helpful.
The takeaway: the feared 'two-faced' behavior is inverted, and unrealistic eval environments are a major culprit.
More from AGI Musings
- Raspbian creator laments $300 Raspberry Pi as AI demand prices out young programmers — evilsocket · 2026-09-10
- RSI is here, just disaggregated: DeepSeek using LLMs to design algorithms — teortaxesTex · 2026-09-10
- RIP Data Scientists: One Prompt Now Does a Week of DS Work in 13 Minutes — mdancho84 · 2026-09-10
- Which Government Jobs Are High-Leverage for AI Safety? Legislation Is the Bottleneck — ShakeelHashim · 2026-09-10
- AI Safety Talent Debate: Frontier Labs vs Government as the Point of Maximum Leverage — ShakeelHashim · 2026-09-10
- AI Pause advocates have no criteria for resuming—so do they actually mean Stop? — menhguin · 2026-09-10