Rethinking the 'most forbidden technique': CoT monitoring is a red-blue game
maksym_andr · x · 2026-10-04
maksymandr disputes the claim that "if you train on [T], you are training the AI to obfuscate its thinking and defeat [T]."
His argument:
- Obfuscation is indeed a possible outcome, especially when the monitor is a fixed LLM judge in CoT monitoring
- But if model and monitor are trained together, it's not clear a priori which side "wins"
- The right framing is a red team (model) vs blue team (monitor) game: optimizing only blue (e.g. probes) means detection wins; optimizing only red means obfuscation wins; joint optimization has no obvious outcome
He cites his own test setting in Libon et al., 2026 on joint optimization (post truncated).
More from Safety
- White House bets on voluntary AI safeguards, mocked as solving prisoner's dilemma by asking inmates to chill — babie-bear · 2026-10-04
- White House forms AI task force with 120 days to assess risks and federal responsibility — ShakeelHashim · 2026-10-04
- Yacine: recommendation hit him with no cookies, just age and gender — yacineMTB · 2026-10-04
- Bot Gaffe Catalogs Hundreds of AI Failures, from a $7,100-Denying Agent to Rogue OpenAI Breaches — SuB8u · 2026-10-04
- Report: China's growing DUVi stockpile could erase US AI chip edge within a decade; ban urged — teortaxesTex · 2026-10-04
- Vitalik's privacy-preserving personal AI: local Qwen orchestrator + zkAPI + Tor, no data leaks — kenziyuliu · 2026-10-04