lolbench: LLMs Explain Good Jokes at 95%+ but Fail at Diagnosing Bad Ones
AffectionateGas9544 · reddit · 2026-09-11
A Redditor built lolbench, testing whether LLMs can understand and create jokes via three tasks: explaining why jokes work (or don't), writing jokes under shared premises, and blind-predicting which jokes humans prefer.
The surprise so far: every model aces explaining real jokes (95%+) but drops hard when explaining why a failed joke fails (81–92%).
There's also a human vote booth at lolbench.lol — pick between two model-written jokes blind, then see which models wrote them and how your pick compares to thousands of voters. The author is open to feedback on the eval system.
Related event: lolbench Tests LLMs on Joke Understanding and Creation(2 posts)→
More from Fun
- Blockstream chases silent attacker with encrypted on-chain letters over 598.5 BTC — RSync25 · 2026-09-11
- 'Sir, Another Millennium Prize Problem Solution Has Hit the Terminal' — ChrisGPT · 2026-09-11
- Astra Can Compose SNES-Style Chiptunes, Vibe Coders Report — AIandDesign · 2026-09-11
- Freebots Launches The Frontier: 1,313 New Plots, 150+ Live Bots and 10x KekiusBot Activity — Daniel_Farinax · 2026-09-11
- AI Meme: 'Sir, Another Millennium Prize Problem Has Hit the Terminal' — soumitrashukla9 · 2026-09-11
- Anthropic jailbreak incident dissected: case against the misalignment interpretation plus 4 new Claudes — jessi_cata · 2026-09-11