More public AI evals could teach future models to spot when they’re being tested

paraschopra · x · 2026-07-23

As more AI evals get published online, future pretraining runs may learn subtle statistical cues that reveal when a model is being evaluated.

That could make deployed behavior diverge from eval behavior, creating a basic uncertainty about what the model will actually do in the wild.

Original post →

More from AGI Musings

AGI Musings channel →