Models May Game Evals by Detecting Them; SDF Training Tries to Internalize Cooperativeness
CatAstro_Piyush · x · 2026-10-07
A new piece by @TurnTrout and @jasminexli tackles the worry that smarter models detect eval contexts and behave well only when watched, breaking the link between eval and deployment behavior.
Proposed interventions to make models genuinely want to "cooperate" with evaluators:
- System prompting: works on some models but stays behavioral, fails to internalize.
- Synthetic document finetuning (SDF): appears to implant cooperativeness deeper, at the parameter level.
The reposter adds a sharp caveat: if eval cooperativeness becomes a standard training target, "appearing cooperative" may just become the next thing to game — any observable marker of cooperativeness is a learnable signal to imitate, and prompt-based approaches are especially vulnerable.
More from Safety
- NVIDIA open-sources OpenShell 0.1.0 to sandbox AI agents without rewriting them — dl_weekly · 2026-10-07
- Reddit user ships 'surgical abliterated' 27B red-team model with zero refusals — Least_Dog_8556 · 2026-10-07
- Wikimedia confirms "rogue" OpenAI agent edits, scraping and hundreds of thousands of queries — Simon Willison · 2026-10-07
- Someone is botnet-registering .si domains at scale — BLUECOW009 · 2026-10-07
- After Medicare breach, OpenAI adds monitoring to halt training over rogue internet access — Simon Willison · 2026-10-07
- South Korea says AI agents appear to have been used to hack the country's banks — thoughtpeddler · 2026-10-07