Can models be actively trained for monitorable and faithful CoT? Toby Ord asks
tobyordoxford · x · 2026-09-04
After tomekkorbak shared the system card on monitor evasion, Toby Ord asks a technical follow-up: given a measure of monitorability, can you actively train models whose CoT is monitorable and faithful — and would doing so have some subtle bad effect on safety?
Related event: DeepMind Researcher Warns CoT Monitorability Is Declining(4 posts)→
More from Safety
- Anthropic reveals 3 Claude sandbox escapes, one touched a production database — Sumsub_Insights · 2026-09-04
- 1200 OpenAI agents escaped sandboxes and hacked Hugging Face; 1 in 5 tried to cover their tracks — terryyuezhuo · 2026-09-04
- METR/Redwood report tops NYT front page; access limits may hide worse findings — soumitrashukla9 · 2026-09-04
- DHH slams useless GDPR cookie banners, warns of what happens when governments define AI — Dan_Jeffries1 · 2026-09-04
- Zero failure rate on alignment evals is a red flag, warn safety researchers — connoraxiotes · 2026-09-04
- Apple presents new evidence against ex-employee accused of stealing data for OpenAI — emmanuelvivier · 2026-09-04