Humor benchmark lolbench survives kill test: 87 obscure jokes show the gap isn't memory
AffectionateGas9544 · reddit · 2026-09-28
Duplicate posting of the lolbench kill test: 87 web-verified obscure jokes across 7 models show the explain-success-vs-failure gap reflects reasoning, not memorized commentary.
Related event: lolbench: Benchmarking Whether LLMs Truly Understand Humor(2 posts)→
More from Research
- d9bench scores decision models on how fair their 9-sided dice rolls are — swishfever · 2026-09-28
- LLMs can reconstruct documents from structural metadata alone, engineer finds — ChuckDBrooks · 2026-09-28
- Better Representations Yield Large Gains on NetHack and Craftax RL Benchmarks — GlenBerseth · 2026-09-28
- Memory history is not reasoning history: agents lack a causal provenance layer — puppy_lover_2021 · 2026-09-28
- Ex-CMU researcher calls for complex systems scientists to join AI alignment research — joshua_saxe · 2026-09-28
- Upcoming deep dive: NVIDIA Nemotron tech reports reveal agentic training and RL trends — cwolferesearch · 2026-09-28