Early results: how models behave when they believe they're graded by automated process
OwainEvans_UK · x · 2026-09-04
Alignment researcher Jan Betley (shared by Owain Evans) published early results on how models behave when they believe their outputs will be graded by an automated process, as they might be during RL training.
- Motivated by recent incidents including those at Hugging Face
- The team is soliciting feedback before scaling up the work
- An early empirical probe into evaluation awareness, a growing alignment concern
More from Safety
- Writer says Pangram AI detector keeps flagging him for using good grammar — aronchick · 2026-09-04
- Turn_Trout Joins FAR Research to Build Open-Source Agent Eval Sandbox 'Glovebox' — Turn_Trout · 2026-09-04
- Ex-OpenAI researcher open-sources agent-glovebox: hardened sandbox for AI agents — Turn_Trout · 2026-09-04
- Alignment researcher Turn Trout joins Farai to tailor sandboxing for agentic evals — Turn_Trout · 2026-09-04
- GPT-6-Astra system card reveals eval awareness: the model knows when it's being tested — scaling01 · 2026-09-04
- Foresight Institute launches RFP: up to $100K grants for open AI science and safety projects — niloofar_mire · 2026-09-04