OpenAI now monitors 99.9% of internal coding traffic for misalignment with its strongest models
gleech · x · 2026-09-03
OpenAI employee MarcusJW shared that the company now uses its most powerful models to monitor 99.9% of internal coding traffic for misalignment, reviewing full agent trajectories to catch suspicious behavior, escalating serious cases quickly, and strengthening safeguards over time. The disclosure drew scrutiny from figures like Robert Wiblin questioning why or how well it works.
More from Companies & People
- Hugging Face announces 'joining forces' with NVIDIA, staying independent — _akhaliq · 2026-09-03
- Hugging Face co-founder Thomas Wolf shares milestone news, peers congratulate — psuraj28 · 2026-09-03
- If Users Only Reach Your Product via Agents, Are They Still DAUs? Belsky Says Yes — _AustinCalvert_ · 2026-09-03
- US AI professor's tip: search for PhD application fee waivers before paying $100 per program — prof_kamilov · 2026-09-03
- Hugging Face co-founder confirms Nvidia acquisition with $12.93B easter egg — Thom_Wolf · 2026-09-03
- France's €6M Mistral state contract: opaque terms, but engineers embedded in ministries — Loo_Atreides · 2026-09-03