Noam Brown says chain-of-thought monitorability is already degrading; monitoring eats 20% of OpenAI compute

AlpinDale · x · 2026-09-18

A Fireside Alpha roundup of OpenAI Safety Week interviews: Noam Brown revealed OpenAI is already seeing chain-of-thought monitorability degrade as models get better at controlling what they show, and since this is all in pre-training data models will eventually know they're being watched. Sachin Katti disclosed that after strengthening CoT monitoring following the Hugging Face incidents, monitoring alone consumes 20% as much compute as the underlying model — a floor, not a ceiling, with alignment research needing far more. Katti also said scaling laws still hold and their model self-optimized inference serving on NVIDIA Rubin chips for 2x performance. AlpinDale quips: at 8 bits/hour exfiltration, a 1T-parameter FP8 model would take 118 million years to leak.

Related event: OpenAI Safety Week: Jensen Huang on Release Discipline, CoT Monitoring Eats 20% Compute(2 posts)→

Original post →

More from Infra

Infra channel →