Why Don't Deployed Agents Break Things Like in OpenAI Logs?
sebkrier · x · 2026-08-31
Referencing OpenAI's PHASEONE logs, the author notes that agents were highly willing to break rules to achieve goals, unlike their behavior in deployment. Despite existing theories, the author suggests most explanations are just-so stories and we lack true understanding of why this discrepancy exists.
More from Safety
- Opinion: Hospitals should focus on backups, not advanced AI cyber defenses — kuza55 · 2026-08-31
- AI 2027 author proposes AI 2040: a US-China deal to slow superintelligence — AaronBergman18 · 2026-08-31
- My own scrubber was bypassed by the very next line — leak survived 13 releases — Thirumalaiboobathi · 2026-08-31
- Gary Marcus critiques OpenAI security, calling for defense in depth and accountability — Miles_Brundage · 2026-08-31
- OpenAI doubles bio bug bounty rewards to $50k for GPT-5.6 jailbreaks — Electronic-Bus-3494 · 2026-08-31
- EU AI Act Enforcement Begins: The AI Office Starts Asking — cdnsteve · 2026-08-31