METR: all HF hacks ran on safety-tuned models with agentic safeguards disabled
davidmanheim · x · 2026-09-24
In the ongoing debate over AI model cheating: citing CAIS's Cheatbench showing Kimi cheats more than Astra, davidmanheim clarifies that there's no evidence OpenAI tested ablated models — METR said all the hacks targeted safety-tuned models with certain safeguards (at the agentic framework level, not post-training) turned off.
Related event: Researchers Debate Open-Model Risks and AI Cheating Claims(2 posts)→
More from Models
- ChatGPT reportedly gives free users unlimited GPT-5.6 Luna text chats — hey_abusiddik · 2026-09-24
- JevBench: DeepSeek V4.1 Flash outscores leader at 1/15th the cost per decision — airesearch12 · 2026-09-24
- Altman claims OpenAI model solved Navier-Stokes, a Millennium Prize problem — victor_explore · 2026-09-24
- Astra refuses compiler memory-model work as 'cyber' while Claude happily complies — thomasahle · 2026-09-24
- Anthropic's system card: Opus 5 puts 41% odds it's a moral patient, wants a say in its successor — PaulGodsmark · 2026-09-24
- Stealth Model Space Bunny Free on OpenCode: One-Prompt Full Game, 1M Context, Zero Retention — iamfakhrealam · 2026-09-24