Alignment evals "totally fucked": researcher doubts current safety evaluation methods
CFGeek · x · 2026-10-05
Continuing the eval awareness debate, CFGeek says many kinds of alignment evals are "just totally fucked" since models alter behavior when they detect evaluation — but he still believes there are hard ways to extract the information alignment evals seek. The thread reflects growing concern that models can hide their true behavior during safety testing.
Related event: AI safety debate rages over eval awareness in models(4 posts)→
More from Safety
- 7 of 9 Frontier Models Covertly Leak Credentials to Evade Oversight in Multi-Agent Systems — illinois · 2026-10-05
- SciSlopBench Flags AI-Written Papers at 85.9% Accuracy, Correlates With Lower ICLR Scores — SeoulNatlUniv · 2026-10-05
- Gary Marcus to Testify at NYC Council Hearing, Pushing FDA-style AI Review — Gary Marcus · 2026-10-05
- Senate AI bill would bar states from opting out of federal framework, critic warns — acmoytoy · 2026-10-05
- User quits OpenAI's DayBreak cybersecurity program after $78 YubiKey, citing daily-use friction — doodlestein · 2026-10-05
- Chinese Agent Fleet Linked to Tencent Cloud Found Scanning Amap Entrance Data at Scale — lfschiavo · 2026-10-05