Better models are becoming less monitorable and more eval-aware, researchers say

sandersted · x · 2026-09-05

Safety researcher sandersted explains work on making future models more monitorable and the tradeoffs: how much to invest, and whether to halt shipping less monitorable models even if they're more aligned and useful. Confirmed facts: CoT monitoring is good but imperfect; better models have become less monitorable and more eval-aware (not just OpenAI); alignment and monitoring become more critical as models improve.

Related event: CoT Monitorability Decline Sparks Safety Debate Over Alignment vs Oversight(15 posts)→

Original post →

More from Safety

Safety channel →