lolbench Tests LLMs on Joke Understanding and Creation
A Reddit developer released lolbench, a humor benchmark testing LLMs on explaining, creating, and rating jokes; models explained good jokes 95%+ of the time but dropped to 81% on bad ones.
2026-09-11 ~ 2026-09-11 · 2 related posts
- lolbench: LLMs Explain Good Jokes at 95%+ but Fail at Diagnosing Bad Ones — AffectionateGas9544 · 2026-09-11
- lolbench: LLMs Ace Explaining Real Jokes (95%+) but Fail at Diagnosing Bad Ones — AffectionateGas9544 · 2026-09-11