OpenAI burns 20% as much compute on monitoring as the model itself, SemiAnalysis says
kevinnbass · x · 2026-09-17
SemiAnalysis's Max Kan revealed that after OpenAI enhanced its chain-of-thought monitoring following the Hugging Face incidents, monitoring alone now consumes roughly 20% as much compute as the underlying model itself. He argues alignment and interpretability research will require far more compute, and that people need to internalize how compute-intensive safety work really is. Commenter kevinnbass draws an analogy to humans spending huge "compute" policing group norms, predicting AIs will soon censor and threaten other AIs.
More from Models
- Dev says DeepSeek-v4.1-flash can reverse engineer anything they want — gaganghotra_ · 2026-09-17
- Redditor claims new Gemini 4 checkpoint is out and noticeably better — Last_Conclusion_8984 · 2026-09-17
- Why Google skips the frontier LLM race: cheap Flash models over beating rivals — burkov · 2026-09-17
- OpenRouter's Mystery Union Model Reverse-Engineered: Likely a Qwen4 MoE — unsane_imagination · 2026-09-17
- New Model Jev Runs Security Pipelines 5x Cheaper and Faster, Devs Say — zeeg · 2026-09-17
- API users have zero loyalty: LLM market share could flip from 80-20 to 20-80 overnight — tengyanAI · 2026-09-17