Humor benchmark lolbench survives kill test: 87 obscure jokes show the gap isn't memory
AffectionateGas9544 · reddit · 2026-09-28
The author of lolbench, a benchmark testing whether LLMs truly understand humor, faced a sharp critique: models ace explaining why real jokes work (95%+) but score lower on failed jokes (81-92%) — possibly because famous jokes ship with online commentary, making success explanations mere recall.
- The critique: the gap might measure the distance between remembering and thinking, not reasoning
- The kill test: 87 web-verified obscure jokes with zero analysis anywhere, 7 models, 2 judges, 885 graded pairs, $0
- Result: scores didn't collapse — obscure working jokes score the same as famous ones (mean gap +0.1, all models within confidence interval), and the working-vs-failed gap survives between equally obscure items
The benchmark's axis holds: the gap reflects genuine understanding, not familiarity. Per-model table at lolbench.lol/kill-test.
Related event: lolbench: Benchmarking Whether LLMs Truly Understand Humor(2 posts)→
More from Research
- CMU's DeformX trains robots to whip ropes in sim — UR5e knocks apple off a head with 0cm error — DJiafei · 2026-09-28
- Stanford: self-organizing agent teams beat oracle router by 13.4 points on AIME — Justgototheeffinmoon · 2026-09-28
- Wei Xu shares talks on multilingual LLMs and rollout diversity for GRPO-style RL — cocoweixu · 2026-09-28
- Sakana AI unveils SAIL, scaling in-context imitation learning for robots without retraining — SakanaAILabs · 2026-09-28
- Team completes Lean formalization of Poincaré conjecture proof in 4.7M lines — latticecut · 2026-09-28
- Erik Hoel: Consciousness is a mystery, but its definition isn't — experts mostly agree — erikphoel · 2026-09-28