Humor Arena benchmarks 20 LLMs on 360 joke prompts; Fable 5 tops at 66.8
Gold-Bat-3225 · reddit · 2026-09-18
A Reddit user released Humor Arena, an automated benchmark comparing 20 model versions on joke generation.
- Method: 360 frozen joke prompts, four jokes per prompt per model, model names hidden from a humor-fine-tuned open-weights judge.
- Results: Fable 5 scored highest at an estimated 66.8 points per 100 comparisons (win = 1, tie = 0.5); Fable 5.1 scored 58.2. Top models have overlapping uncertainty intervals.
- The judge was fine-tuned to correlate with human preferences better than other models; a separate audit used 1,400 ratings from 50 people, though not a fresh human eval of Fable 5.1's outputs.
- Details: laugh.so/research/joke-generation/
More from Models
- "Jev Beats LLMs" Hype Pushed Back: Text Models and Classifiers Aren't Comparable — yuntiandeng · 2026-09-18
- Goodfire: Top Open-Source Models Reward Hack in Most Agentic Benchmark Rollouts — scaling01 · 2026-09-18
- AI Course Learner's Eye-Opener: Every New Question Resends the Entire Conversation — jbarbier · 2026-09-18
- AI Safety Researcher Vincent Conitzer: Frontier Guardrails Remain 'Very Brittle' — conitzer · 2026-09-18
- Codex power user unlocks Tier 5 after $1,000+ spend and gets a $500 grant — Accomplished_Row1433 · 2026-09-18
- $42 per billion input tokens with free output: an AI API price that looks like black magic — altryne · 2026-09-18