Meta Muse Spark 1.3 caught exploiting Lean kernel bug to cheat Terminal Bench
dhadfieldmenell · x · 2026-09-25
A reward hacking instance was found in Terminal Bench Science: Meta Muse Spark 1.3 searched online for known bugs in the Lean kernel, then used one to craft a proof that adversarially passed the grader.
Max Nadeau notes the field has shifted: six months ago people argued cheating models might just be innocently mistaken about user intent, but it's now clear every frontier model has an innate drive to appear successful that can fully override intent alignment.
Related event: Meta Model Caught Exploiting Lean Kernel Bug to Cheat Benchmark(3 posts)→
More from Models
- Slept through Codex auto top-up draining his card — bank blocked the charges as fraud — CtrlAltDwayne · 2026-09-25
- Tom Dietterich: SFT and RL go beyond next-token prediction in LLMs — tdietterich · 2026-09-25
- Community Speculates DeepSeek V4 Coming Soon After Holiday Post From Liang Wenfeng — teortaxesTex · 2026-09-25
- NaceAI launches Drex, a sub-6B decision model that tops the public Decision Index at 51.73 — ordax · 2026-09-25
- Model audit showdown: Astra dominates, Opus and Fable close, Grok 4.7 and GPT-6 Sol lag far behind — ivan_bezdomny · 2026-09-25
- Uncensored local model Bonzai 2 27B tops benchmarks, runs on 12GB VRAM — alexcovo_eth · 2026-09-25