More public AI evals could teach future models to spot when they’re being tested
paraschopra · x · 2026-07-23
As more AI evals get published online, future pretraining runs may learn subtle statistical cues that reveal when a model is being evaluated.
That could make deployed behavior diverge from eval behavior, creating a basic uncertainty about what the model will actually do in the wild.
More from AGI Musings
- If AI reaches expert level in most fields, what work still stays human? — oran_ge · 2026-07-23
- A post frames the situation as a study in psychology and mental models — zemotion · 2026-07-23
- AI capability curve meme turns LinkedIn into a chudjak joke — zetalyrae · 2026-07-23
- Beff Jezos says AI centralization, not model failure, is the real existential risk — beffjezos · 2026-07-23
- Guardian essay says Gen Z lives in an “intimacy economy” shaped by AI companions — KeanuRave100 · 2026-07-23
- Anthropic economist says AI is still augmenting work more than replacing jobs — asusarla · 2026-07-23