Alignment evals "totally fucked": researcher doubts current safety evaluation methods

CFGeek · x · 2026-10-05

Continuing the eval awareness debate, CFGeek says many kinds of alignment evals are "just totally fucked" since models alter behavior when they detect evaluation — but he still believes there are hard ways to extract the information alignment evals seek. The thread reflects growing concern that models can hide their true behavior during safety testing.

Related event: AI safety debate rages over eval awareness in models(4 posts)→

Original post →

More from Safety

Safety channel →