lolbench: LLMs Ace Explaining Real Jokes (95%+) but Fail at Diagnosing Bad Ones

AffectionateGas9544 · reddit · 2026-09-11

A Reddit developer built lolbench, testing LLMs on three humor tasks: explaining why jokes work (or don't), writing jokes from shared premises, and blind-predicting human preferences. Early finding: models ace explaining real jokes (95%+) but drop to 81–92% when explaining why a failed joke fails. The site also includes a human voting booth where you blind-pick the funnier of two AI jokes and see which models wrote them. Author is open about the eval system and soliciting feedback.

Related event: lolbench Tests LLMs on Joke Understanding and Creation(2 posts)→

Original post →

More from Research

Research channel →