lolbench: LLMs Explain Good Jokes at 95%+ but Fail at Diagnosing Bad Ones

AffectionateGas9544 · reddit · 2026-09-11

A Redditor built lolbench, testing whether LLMs can understand and create jokes via three tasks: explaining why jokes work (or don't), writing jokes under shared premises, and blind-predicting which jokes humans prefer.

The surprise so far: every model aces explaining real jokes (95%+) but drops hard when explaining why a failed joke fails (81–92%).

There's also a human vote booth at lolbench.lol — pick between two model-written jokes blind, then see which models wrote them and how your pick compares to thousands of voters. The author is open to feedback on the eval system.

Related event: lolbench Tests LLMs on Joke Understanding and Creation(2 posts)→

Original post →

More from Fun

Fun channel →