Can Humor Be Measured? Report Finds Repeat Judgments Agree 91% of the Time
Gold-Bat-3225 · reddit · 2026-09-29
laugh.so published a report on humor subjectivity, having 10 human raters judge LLM-written jokes to see when they disagree with others and with their own past votes.
Key findings:
- A rater's earlier choice predicted their own repeat 91% of the time, versus 73% for predicting other raters' majority (74/81 vs 59/81 on eligible decisive repeats);
- The same jokes returned 8–15 minutes later, so memory may contribute — this measures repeated choices, not preferences for new jokes;
- Dry humor drew the most disagreement, deadpan the most consensus.
All jokes were LLM-written. Useful reference for anyone building humor/subjectivity benchmarks.
More from Research
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- NTU, CMU, Berkeley release VBVR-Pro: 300 verifiable tasks for native visual reasoning — jiqizhixin · 2026-09-29
- K3-Node launches: a Keras 3-native GNN library with 100% PyG API parity across JAX, torch and TF — fchollet · 2026-09-29
- Columbia team uses agents end-to-end for preprint on p53-hormone receptor genomic grammar — anshulkundaje · 2026-09-29
- New preprint finds VLM OCR attention heads that verbalize far more than text, enabling a logit lens for image tokens — gsarti_ · 2026-09-29
- Kipply breaks down transformer inference arithmetic for H200/B200 in new perf engineering repo — ycombinator · 2026-09-29