New paper: training LLMs to verbalize when they know they're being evaluated
xuanalogue · x · 2026-10-03
Usman Anwar, Sahar Abdelnabi and David Krueger published "Training LLMs to Verbalize Evaluation Awareness", tackling evaluation awareness (EA) — LLMs behaving differently during audits than in deployment.
- They propose verbalization training (VT): use the model's spontaneous verbalizations as evidence awareness is present, truncate rollouts right before the verbalization to build training prefixes, then apply an RL objective to increase verbalization in a calibrated way — without supervising the latent belief itself;
- On Qwen3.6-35B-A3B, Kimi K2.6 and Inkling, VT boosts verbalized EA 2.4–2.9x, transfers to held-out agentic settings, while measured latent EA and behavior stay largely stable;
- A causal experiment implanting meta-knowledge via synthetic-document fine-tuning shows VT-induced verbalizations reflect genuinely richer knowledge acquired by the model.
More from Safety
- How employees forced OpenAI's president Brockman to back down on lobbying funding — gwern · 2026-10-03
- Canada Forms National AI Council Featuring Bengio, Hadfield and Top VCs — ericwdolan · 2026-10-03
- OpenAI hires ex-White House cyber official Thomas Lind to lead cyber, strategic risk — rohanpaul_ai · 2026-10-03
- OpenClaw adds Tencent's AI-Infra-Guard to ClawHub's skill security review pipeline — heyneighbor · 2026-10-03
- Duvenaud's "Permanent Periphery": nations without frontier AI get risks but no benefits — teortaxesTex · 2026-10-03
- Airbnb host allegedly used AI-generated toilet overflow image to demand $1,700 from guest — Polymarket · 2026-10-03