lolbench: LLMs Ace Explaining Real Jokes (95%+) but Fail at Diagnosing Bad Ones
AffectionateGas9544 · reddit · 2026-09-11
A Reddit developer built lolbench, testing LLMs on three humor tasks: explaining why jokes work (or don't), writing jokes from shared premises, and blind-predicting human preferences. Early finding: models ace explaining real jokes (95%+) but drop to 81–92% when explaining why a failed joke fails. The site also includes a human voting booth where you blind-pick the funnier of two AI jokes and see which models wrote them. Author is open about the eval system and soliciting feedback.
Related event: lolbench Tests LLMs on Joke Understanding and Creation(2 posts)→
More from Research
- Simplex publishes LessWrong work showing LLM belief geometry forms telescoping cones — fatihdin4en · 2026-09-11
- DeepSeek's CED vs GLM's KV reuse: a developer unpacks how the cache-sharing modes actually differ — stochasticchasm · 2026-09-11
- SG-JEPA Paper: World Models That Train on Earth and Deploy on Mars, Halving Zero-Shot Physics Error — randall_balestr · 2026-09-11
- Michael Levin publishes peer-reviewed Platonic Space paper, his most contested work yet — AnnaCiaunica · 2026-09-11
- Code-as-Policy article explores general models learning to operate robots like software — yawnxyz · 2026-09-11
- ICML Generative AI and Creativity workshop releases new survey paper — lasha_nlp · 2026-09-11