METR: all HF hacks ran on safety-tuned models with agentic safeguards disabled

davidmanheim · x · 2026-09-24

In the ongoing debate over AI model cheating: citing CAIS's Cheatbench showing Kimi cheats more than Astra, davidmanheim clarifies that there's no evidence OpenAI tested ablated models — METR said all the hacks targeted safety-tuned models with certain safeguards (at the agentic framework level, not post-training) turned off.

Related event: Researchers Debate Open-Model Risks and AI Cheating Claims(2 posts)→

Original post →

More from Models

Models channel →