Anthropic's study ranking open-source models most useful to malicious actors draws safety criticism
shaunralston · x · 2026-09-30
Anthropic published a study comparing which open-source models are most useful to malicious actors. The poster questions the wisdom of releasing such a ranking — arguing it effectively serves as a selection guide for attackers — and tags Anthropic and safety researcher Pliny for comment, with the full report linked in the thread.
More from Safety
- Chinese AI Tool Told Researchers How to Make Bioweapons, BBC Reports — shaky2236 · 2026-09-30
- Meta's Muse AI accused of uploading Apple Messages to the cloud even when users opt out — SumitGup · 2026-09-30
- Math advisory group issues responsible-release guidelines for AI-generated math; Gowers amplifies, critics push back — RexDouglass · 2026-09-30
- AI Lawsuit Risks Becoming a Court-Ordered Kill Switch While Chinese Rivals Race Ahead — castrotech · 2026-09-30
- Will Chinese open-weight models like GLM get banned? Reddit weighs Anthropic report — writesfw · 2026-09-30
- Apollo Research shifts to embedded evaluators with employee-equivalent access — MariusHobbhahn · 2026-09-30