Zvi on Anthropic's misuse report: seven harm areas, and distillation deserves the list spot
TheZvi · x · 2026-09-15
Zvi reviews Anthropic's latest threat report covering disrupted misuse activity from December 2025 to August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation.
His core take: many bad actors try to misuse Claude and mostly fail—or at least Anthropic believes so. If the report shares the worst cases or anything close, closed models are looking better on the misuse front than he expected.
He pushes back on the view that distillation is a trivial inclusion: as the report itself notes, distillation boosts performance across nearly every task, enabling the other six harm categories.
More from Models
- Sam Altman teases 'big ship this week' for OpenAI, more at DevDay — altryne · 2026-09-15
- Users report being routed to GPT-6 Sol, early impressions consistently positive — kimmonismus · 2026-09-15
- Claude Fable 5.1 holds #1: 1M-token input, ties GPT 6 Astra on new benchmarks — DeepLearningAI · 2026-09-15
- Tests show OpenAI hasn't changed Codex quotas: 800M+ Astra tokens per cycle — daniel_mac8 · 2026-09-15
- AlphaSense tests Fable and Astra as financial researchers: top score but only on par with Opus-5 — CShorten30 · 2026-09-15
- Grok 4.7 may drop today from xAI, unless delays strike again — mark_k · 2026-09-15