62-person blind test of 48 LLM jokes shows models are getting funnier, Astra tops with 50%+ laughs
paraschopra · x · 2026-10-06
Paras Chopra ran an experiment testing whether frontier models are getting better at comedy: 62 respondents rated 48 jokes from six models.
Humor did improve over time — Astra came out on top, earning chuckles or laughter from over 50% of raters. The study probes whether capability gains in verifiable domains like math and coding transfer to soft, hard-to-verify skills like joke writing.
Full methodology and results are on his Substack.
Related event: Blind Test of 48 LLM Jokes Shows AI Humor Is Improving(2 posts)→
More from Fun
- Claude spends 18 hours solo-building a Windows XP clone called Macrohard — FinanceYF5 · 2026-10-06
- Opus 5.5 Builds a Fully Animated Pixel-Art Scene in TypeScript in ~10 Minutes — iamfakhrealam · 2026-10-06
- DINOv3 + SAM Local Vision Pipeline Counts Surgical Instruments in Real Time — MaziyarPanahi · 2026-10-06
- Claude Opus 5.5 took 90,000 screenshots of its own game over two days to polish the visuals — prasenx · 2026-10-06
- StealthGPT claims it killed AI detector Pangram, sparking fraud backlash — paulnovosad · 2026-10-06
- An AI-composed indie-pop song about indium bubbles, made on GPUs — davidmanheim · 2026-10-06