Humor Benchmark: Gemini 3.7 Wins, GPT-4o Struggles to Be Funny

scaling01 · x · 2026-08-30

Laugh Labs released a report benchmarking the sense of humor of 16 top LLMs against >100k human ratings.

Key Findings:

Original post →

More from Models

Models channel →