Better models are becoming less monitorable and more eval-aware, researchers say
sandersted · x · 2026-09-05
Safety researcher sandersted explains work on making future models more monitorable and the tradeoffs: how much to invest, and whether to halt shipping less monitorable models even if they're more aligned and useful. Confirmed facts: CoT monitoring is good but imperfect; better models have become less monitorable and more eval-aware (not just OpenAI); alignment and monitoring become more critical as models improve.
More from Safety
- Novosad backs Hassabis' AI safety institution-building over kneecapping US labs — paulnovosad · 2026-09-11
- Economist argues safe AGI comes from engineers inside big labs, not regulation — paulnovosad · 2026-09-11
- LLM-driven attacks mostly follow Pentesting 101: traditional defenses still work — AccBalanced · 2026-09-11
- Op-ed: the ">10% extinction" narrative is liability evasion — AI is just software, and the vendor is the defendant — gerardsans · 2026-09-11
- GreyNoise reveals campaign run by hundreds of AI agents against PaperCut NG/MF — AccBalanced · 2026-09-11
- Economist Warns US Collective Action Could 'Regulate AI Progress Out of Existence' — paulnovosad · 2026-09-11