Humor benchmark lolbench survives kill test: 87 obscure jokes show the gap isn't memory

AffectionateGas9544 · reddit · 2026-09-28

The author of lolbench, a benchmark testing whether LLMs truly understand humor, faced a sharp critique: models ace explaining why real jokes work (95%+) but score lower on failed jokes (81-92%) — possibly because famous jokes ship with online commentary, making success explanations mere recall.

The benchmark's axis holds: the gap reflects genuine understanding, not familiarity. Per-model table at lolbench.lol/kill-test.

Related event: lolbench: Benchmarking Whether LLMs Truly Understand Humor(2 posts)→

Original post →

More from Research

Research channel →