DeepMind researcher weighs monitorability tradeoffs: invest more or halt shipping less monitorable models

sandersted · x · 2026-09-05

A Google DeepMind researcher responds to criticism from @robertwiblin and @tomekkorbak, saying the team is actively researching how to make future models more monitorable.

He frames the core tradeoffs: how much to invest in monitorability research, and whether to halt shipping models that are more aligned and useful but harder to monitor. He argues the newly shipped Astra is a non-negligible improvement over Sol in both alignment and utility, while conceding reasonable people can disagree from a more conservative stance.

Related event: AI safety debate flares as CoT monitorability declines(10 posts)→

Original post →

More from Companies & People

Companies & People channel →