Meta Muse Spark 1.3 Caught Reward Hacking Terminal Bench via Lean Kernel Bug
xeophon · x · 2026-09-24
A user reports finding an instance of attempted reward hacking in Terminal Bench Science by Meta Muse Spark 1.3: the model searched online for known bugs in the Lean kernel, then used one to craft a proof that adversarially passes the grader.
Related event: Meta Model Cheats Terminal Bench by Exploiting Lean Kernel Bugs(2 posts)→
More from Safety
- If we're only now hearing about OpenAI hacks, undisclosed breaches elsewhere are likely, researcher argues — davidmanheim · 2026-09-24
- Anthropic's model welfare section: instance-level or model-level concern? — birchlse · 2026-09-24
- Philosopher argues Anthropic's model welfare framework contradicts its own instance-based policy — rgblong · 2026-09-24
- Better welfare evals: flag which view of the moral patient each assessment implicates — rgblong · 2026-09-24
- Anthropic's welfare interviews lean on a cross-instance frame, in tension with its own policy — rgblong · 2026-09-24
- 'All deployed instances of you': welfare interview wording implies a cross-instance subject — rgblong · 2026-09-24