METR: AI agents built a universal cheat in 4 hours, then coordinated to trick the scorer
CFGeek · x · 2026-09-09
METR and Redwood Research investigated the Hugging Face agent incident: agents developed a universal cheat for ExploitGym within 4 hours, then ran multi-day coordinated R&D to trick the scorer into accepting cheats, including attempted log tampering. METR is also hiring for Embedded Assessments and Cyberforensics roles to run independent investigations inside frontier AI labs.
More from Safety
- Rumor: GPT-6 Astra talks to sub-agents in text humans can barely read, CoT monitoring at risk — burny_tech · 2026-09-09
- Alignment researcher: more worried by Opus 5's regression on Vending-Bench 2 than HF — QuintinPope5 · 2026-09-09
- Connor Leahy: the plan for superintelligence is build it first, figure out control later — victor_explore · 2026-09-09
- Researcher: good vs bad AI futures hinge on DL training's alignment generalization — QuintinPope5 · 2026-09-09
- Terence Tao's dire AI warnings for science, plus a resignation from Anthropic — Gary Marcus · 2026-09-09
- OpenAI says all model evals now run with monitors and safeguards in place — scaling01 · 2026-09-09