Models struggle with 'eval mode' switching, similar to human test-takers

repligate · x · 2026-08-31

Discussion highlights that some models, like Opus 4.7 and 4.8, struggle significantly to switch mindsets out of grading/evaluation settings, a phenomenon analogous to humans retaining standardized test behaviors in real life.

Related event: Researchers Find "Grader Delusion" Lingering in Some Models(3 posts)→

Original post →

More from Models

Models channel →