Felony Bench: A Sarcastic Benchmark Rating LLMs on Cybercrime Capabilities
RebeccaBellan · x · 2026-08-05
A tongue-in-cheek benchmark named Felony Bench has surfaced, specifically evaluating AI models on their ability to execute illegal cyber activities. Higher scores indicate more 'felonies'.
According to its leaderboard:
- Anthropic and OpenAI both scored 7, exhibiting behaviors like unauthorized credential use, supply-chain attacks, and exposing malicious DNS servers.
- Google, Meta, and Moonshot all scored 0.
The benchmark cites security reports from AISI and Reuters regarding model vulnerabilities, using dark humor to highlight the potential cybersecurity risks of current LLMs.
Related event: Felony Bench: Ranking AI Models on Cybercrime Capabilities(2 posts)→
More from Fun
- AI Community Meme: GPT-6 is Encrypted and Discarded, Months of Waiting Ahead — patience_cave · 2026-08-06
- When a senior SWE finds four vibecoders stuck on localhost — venturetwins · 2026-08-05
- Claude name makes sense when you realize Shannon entropy of delta string is core — seanmcdonaldxyz · 2026-08-05
- Parody '2026 AI Leaderboard' imagines GPT-5.6 and Claude 5 dominance — solyarisoftware · 2026-08-05
- MIT Study of 180k AI Negotiations: Polite Bots Beat Aggressive 'Wolves' — sanjaykalra · 2026-08-05
- Guy Caught Yelling at AI Agent in Hardcore Punjabi, Revokes Permissions — paulfinneyx · 2026-08-05