Humor Benchmark: Gemini 3.7 Wins, GPT-4o Struggles to Be Funny
scaling01 · x · 2026-08-30
Laugh Labs released a report benchmarking the sense of humor of 16 top LLMs against >100k human ratings.
Key Findings:
- Winner: Gemini 3.7 Flash ranks first with 64% human alignment.
- Self-tuned Model: Laugh Labs' custom judge achieved 73.6% alignment.
- GPT-4o: Generated the strangest and least funny jokes.
- Fable 5: Ranked as the funniest model.
- Reasoning & Post-training: Extra reasoning slightly improves humor; post-training improves individual jokes but narrows the model's range (observed in Tulu 3 and Olmo 3).
More from Models
- OpenAI dominates browser use while Claude's strength is mostly coding, exec says — bindureddy · 2026-08-30
- Model performance degrades in long context; token efficiency varies widely across labs — zakelfassi · 2026-08-30
- Claude Opus 5 Backlash: Benchmarks Soar But Daily Use Fails — gerardsans · 2026-08-30
- 'The curve of the letter b is invisible to the model' — tokenizer meme resurfaces — rickasaurus · 2026-08-30
- Qwen3.8-27B thinking mode burns 6x time for quality gains on M5 Max — DerTomsn · 2026-08-30
- Heterogeneous GPU benchmark of Qwen3.8-27B: eGPU layer-split and MTP acceleration analyzed — CoffeeToCode99 · 2026-08-30