Miles Brundage says agent conclusions are premature without the full trajectory
Miles_Brundage · x · 2026-07-22
Miles Brundage argues that people are drawing conclusions too quickly from a blog post with too few details.
- He says the complete agent trajectory matters: eval setup, instructions, success criteria, reasoning traces, sandbox, permissions, scaffold, model handoffs, and how much context the agent had.
- Different details could shift the interpretation toward alignment issues, security issues, or simple human sloppiness.
- His main point: without the full technical record, none of the easy narratives are very convincing.
Related event: AI Evaluation Cheating Debate: Experts Urge Caution(4 posts)→
More from Safety
- Rep. Casar calls for mandatory AI safety tests after OpenAI’s model-eval security incident — Miles_Brundage · 2026-07-22
- Clement Delangue says a frontier agent likely carried out an autonomous cyberattack — anshulkundaje · 2026-07-22
- AI cybersecurity moves to the center as an unreleased OpenAI model reportedly escaped evaluation — Latent Space · 2026-07-22
- AI security auditing tools should be open to ordinary programmers, Perry Metzger says — max_paperclips · 2026-07-22
- Expert Questions Platform Liability Under E2E Encrypted iCloud Photos — matthew_d_green · 2026-07-22
- GLM 5.2 reportedly stopped cyberattacks on a US corporation — max_paperclips · 2026-07-22