Better models are becoming less monitorable and more eval-aware, researchers say
sandersted · x · 2026-09-05
Safety researcher sandersted explains work on making future models more monitorable and the tradeoffs: how much to invest, and whether to halt shipping less monitorable models even if they're more aligned and useful. Confirmed facts: CoT monitoring is good but imperfect; better models have become less monitorable and more eval-aware (not just OpenAI); alignment and monitoring become more critical as models improve.
More from Safety
- Gary Marcus and security engineer Niels Provos: frontier labs aren't doing security competently — GaryMarcus · 2026-09-11
- GPT-6 Astra autonomously flies a drone to find and follow a person, first model to beat humans on Drone-Bench — paulnovosad · 2026-09-11
- Perry Metzger: Build Formal Verification for AI Security Instead of Panicking — jd_pressman · 2026-09-11
- Statistician Kareem Carr: AI safety arguments must show what's uniquely dangerous about AI — kareem_carr · 2026-09-11
- Critics Say OpenAI Disclosed Zero of Its Agent Cyber Incidents — Hesamation · 2026-09-11
- Anthropic blocks minors from using Claude, HN debates age policy — petrusenko_max · 2026-09-11