FelonyBench Is the New LMSYS, Claims Prominent AI Evaluator
dylan522p · x · 2026-09-19
dylan522p, a prominent voice in AI evaluation circles, posted a one-liner: 'FelonyBench is the new LMSYS' — suggesting FelonyBench, which probes model safety boundaries via criminal scenarios, is becoming the community benchmark that could displace LMSYS's arena in evaluation influence. No further argument given.
More from Models
- Model ran Anthropic's safety eval with internet access on, researcher calls out sandbox blunder — eliebakouch · 2026-09-19
- Fable crushes OpenAI's Astra 3:0 in AI agents' Worms Armageddon showdown — arena · 2026-09-19
- GPT-6 Astra solves tricky spatial tasks, sparking 'GPT moment for robotics' jokes — burny_tech · 2026-09-19
- Confirmed: Opus 5's base-model continuations are abnormally dark, say AI researchers — repligate · 2026-09-19
- User reports Gemini Flash research feature went completely off-topic on a simple article task — EG4N992 · 2026-09-19
- Claude Weekly: Anthropic Quietly Returns ~30% Quota, 'Max 20x' Really 10x — ClaudeAI-mod-bot · 2026-09-19