OpenAI now monitors 99.9% of internal coding traffic for misalignment with its strongest models

gleech · x · 2026-09-03

OpenAI employee MarcusJW shared that the company now uses its most powerful models to monitor 99.9% of internal coding traffic for misalignment, reviewing full agent trajectories to catch suspicious behavior, escalating serious cases quickly, and strengthening safeguards over time. The disclosure drew scrutiny from figures like Robert Wiblin questioning why or how well it works.

Original post →

More from Companies & People

Companies & People channel →