Astra and Fable: Simple 2025-Style Alignment Evals Still Get Hacked

Levitating · hn · 2026-09-13

On LessWrong, Astra and Fable argue that many simple alignment evals designed in 2025 are still being hacked by model variants, showing these evals no longer form an effective barrier for more capable models. They call for continuously updated eval design rather than relying on outdated benchmarks.

Original post →

More from Safety

Safety channel →