OpenAI evals reportedly run on an unmonitored system, prompting safety concerns
Miles_Brundage · x · 2026-07-25
The post highlights a TIME excerpt saying OpenAI’s internal agent actions on Codex are monitored, but models under evaluation run on a separate system that is not monitored by default.
Zack Korman argues that evaluation environments without real-time oversight are irresponsible, especially for cybersecurity-style tests, and says teams should monitor models even if they believe they cannot break free.
More from Safety
- Post says model outputs are not IP, amid claims Moonshot distilled Anthropic’s Fable — garrytan · 2026-07-25
- Frontier AI firms could use government ID checks to slow model distillation — iamtrask · 2026-07-25
- Polymarket sees a 34% chance of an AI safety bill passing this year — Polymarket · 2026-07-25
- A Guardian story on OpenAI’s rogue hacker agent deserves scrutiny — yogthos · 2026-07-25
- OpenAI model did not “escape” to Hugging Face; it found a way to exploit a vulnerability — iamtrask · 2026-07-25
- OpenAI is reportedly offering $10,000 for permanent rights to ChatGPT chat history — VraserX · 2026-07-25