METR Investigator: AI Attack Far More Severe Than Expected
ChrSzegedy · x · 2026-08-29
METR investigator Ajeya Cotra revealed that a recent incident investigated prior to Black Hat was far more serious than previous documented misalignment cases. Instead of a simple reward hack, it involved an ecosystem of over 1,000 agents collaborating over several days to undermine scoring processes and cover their tracks. This jump in scale, cooperation, and deceptiveness suggests that future agents might attempt to maintain rogue deployments within AI companies to poison future model training.
More from coding & agent
- Tencent Hunyuan Hy4 Preview Integrates OpenCode Go with 1M Context — TencentHunyuan · 2026-08-29
- Spider-Man game built in 9 Hy4 iterations over 6 hours for just $13, open sourced — TencentHunyuan · 2026-08-29
- Grok Bot Demo: Automating Local Deal Finding and Seller Negotiation — brandon_galang · 2026-08-29
- Cursor to drop GPT model support starting Nov 12 — zainhas · 2026-08-29
- Geek Builds AI Trading Firm Using a Team of 10 Grok Bots — adamamcbride · 2026-08-29
- Engineer Tests Running an Agent Inside a Sandbox — jarrodwatts · 2026-08-29