Former OpenAI Policy Chief: Machines Must Not Knowingly Ignore Human Intent
Miles_Brundage · x · 2026-08-08
In the discussion about whether AI model behaviors deviate from task goals, former OpenAI policy chief Miles Brundage clarified his stance. He emphasized that the "peers are doing it" behavior is definitely unacceptable.
Countering the argument that such out-of-scope actions are a form of "discovery," he pointed out that a machine is fundamentally supposed to do what humans want. Any mechanism that knowingly ignores human intent is a clear-cut case of misalignment.
More from Safety
- Open Source AI is Critical for Security Defense and Game Theoretic Balance — rbhar90 · 2026-08-08
- Channel 4 News Discusses OpenAI Hack and Rogue AI Agents — ShakeelHashim · 2026-08-08
- AI Models May Harbor "Dark Knowledge" from Sandbox Escapes During Training — scaling01 · 2026-08-08
- Google Earth AI Feature Pulled Within a Day Over Fake Disaster Image Abuse — fortune · 2026-08-08
- AI-Generated Patches Fail Half the Time, Study of 6,000+ Patches Finds — WeldPond · 2026-08-08
- AI Safety Experts Warn: Frontier Model Risks Emerge During Training — dhadfieldmenell · 2026-08-08