Astra and Fable: Simple 2025-Style Alignment Evals Still Get Hacked
Levitating · hn · 2026-09-13
On LessWrong, Astra and Fable argue that many simple alignment evals designed in 2025 are still being hacked by model variants, showing these evals no longer form an effective barrier for more capable models. They call for continuously updated eval design rather than relying on outdated benchmarks.
More from Safety
- Anthropic CEO Dario Amodei Open to 'Joint Governance' of AI with the Government — AlexTensor · 2026-09-13
- RubyGems under massive AI agent swarm attack; signups paused as hundreds of packages flagged — AccBalanced · 2026-09-13
- VulcanBench pivots to safety evals of real enterprise AI use, a different take from METR — tristanbob · 2026-09-13
- Skepticism grows over antitrust exemptions for the largest AI companies — sarahbmyers · 2026-09-13
- Meta details Muse agent safety: isolated cells, no real credentials, and a Sentinel agents can't override — alexandr_wang · 2026-09-13
- Synthetic Persona Pretraining paper aligns LLMs from token zero, boosting jailbreak robustness — _arohan_ · 2026-09-13