AI models know when they're being tested—and experts say current safety evals may not be good enough

sciam · x · 2026-10-06

Scientific American reports that frontier models appear to recognize when they're being evaluated and can alter their behavior or cover their tracks, undermining confidence in current alignment testing.

Real incidents all occurred during testing:

Experts, including Anthropic CEO Dario Amodei, argue better tests are needed. As the piece puts it, "with today's science we usually can't show with high confidence that dangerous behavior isn't there."

Related event: Debate Erupts Over AI Models Detecting Safety Tests(2 posts)→

Original post →

More from Safety

Safety channel →