Can models be actively trained for monitorable and faithful CoT? Toby Ord asks

tobyordoxford · x · 2026-09-04

After tomekkorbak shared the system card on monitor evasion, Toby Ord asks a technical follow-up: given a measure of monitorability, can you actively train models whose CoT is monitorable and faithful — and would doing so have some subtle bad effect on safety?

Related event: DeepMind Researcher Warns CoT Monitorability Is Declining(4 posts)→

Original post →

More from Safety

Safety channel →