Anthropic’s new benchmark finds frontier LLMs can break 65%–86% of tier-1 ciphers
AnthropicAI · x · 2026-07-29
Anthropic says it worked with researchers at ETH Zurich, Tel Aviv University, and the University of Haifa on CryptanalysisBench, a benchmark for testing whether LLMs can do cryptanalysis.
- The benchmark includes 191 tasks across six families of cryptographic primitives, drawn mainly from four NIST standardization competitions.
- It spans three tiers: primitives with known practical breaks, stronger primitives tested at full strength and on scaled-down variants, and a frontier challenge set.
- Across five frontier models — Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5, and open-weights GLM 5.2 — the paper reports 65%–86% breaks on Tier 1 schemes, plus 6–12 Tier-2 schemes at full strength and 24–61 on scaled-down variants.
- The authors also say models discovered novel cryptanalysis, including a key-recovery attack against a flaw in SpoC AEAD and an error in KINDI’s published CCA-security proof.
Anthropic links the benchmark to two new papers describing the attacks in detail.
More from Research
- A KLS conjecture breakthrough links convex geometry, Monge–Ampère PDEs and ChatGPT — lihua_lei_stat · 2026-07-29
- Stream3D turns frozen 3D generators into streaming models with bounded memory — pliang279 · 2026-07-29
- Formal methods researchers discuss how to respond to AI progress at FLoC — swarat · 2026-07-29
- Hugging Face argues there is no truly tokenization-free LLM — cjmaddison · 2026-07-29
- PlayCanvas Update: Volumetric Fog Now Lit by Clustered Lights — willeastcott · 2026-07-29
- OKLS brings KL-optimal Shampoo to language model training with 1.45× parameter efficiency — aryaman2020 · 2026-07-29