Zvi digests the Claude Fable 5.1 system card: alignment risk up to 'low', prompt injection near solved

Don't Worry About the Vase (Zvi) · rss · 2026-09-05

Zvi Mowshowitz analyzes Anthropic's 200+ page system card for Claude Fable 5.1 — the same model as Mythos 5.1 with classifiers layered on top, and at release the most capable publicly available AI model. It's a substantial but incremental upgrade over Fable 5, with cheaper cache reads and a nicer user experience; comparisons to GPT-6-Astra await more data.

Key safety findings

Zvi also notes CB-2/Autonomy-2 evaluations have drifted from formal tests toward vibe checks, which works only if labs act responsibly — not a basis for robust regulation.

Original post →

More from Models

Models channel →