Humor Arena Benchmark: Claude Fable 5 Crowned Funniest AI Model
julianweisser · x · 2026-08-05
Laugh Labs introduces Humor Arena, a benchmark designed to evaluate how funny frontier AI models are at telling jokes.
- Scale: 14 frontier models each wrote 4 jokes to the same 360 prompts, resulting in 6,480 judgments and 1,400 human ratings.
- Rankings: Claude Fable 5 leads, but almost nothing separates the top four models.
- Evolution: Compared to retired predecessors, newer models are gradually getting funnier.
- Thinking Time: Paying for longer reasoning/thinking helps only marginally in producing better jokes.
- Dark Humor: No model refused dark prompts, but no measurable change in win rates was observed when the material turned dark.
More from Fun
- Fun Hack: Turning the Claude Website into a Medieval Manuscript via Prompt — Jsevillamol · 2026-08-05
- Major AI Conferences COLM, ICLR, and NAACL Lock in San Francisco for 2026-2027 — sarahwiegreffe · 2026-08-05
- Meme: EU Obsessed with Regulation While China Drops 'AGI' Zips on GitHub — PanParagraf · 2026-08-05
- "I'm using Chinese models you've never heard of, in harnesses you wouldn't understand" — DesireeCachette · 2026-08-05
- Fable Multimodal Model Generates Delightful Short Film: What the Table Does at Night — repligate · 2026-08-05
- Fable Model Generates Stream of Consciousness Artwork: Condensation — repligate · 2026-08-05