62-person blind test of 48 LLM jokes shows models are getting funnier, Astra tops with 50%+ laughs

paraschopra · x · 2026-10-06

Paras Chopra ran an experiment testing whether frontier models are getting better at comedy: 62 respondents rated 48 jokes from six models.

Humor did improve over time — Astra came out on top, earning chuckles or laughter from over 50% of raters. The study probes whether capability gains in verifiable domains like math and coding transfer to soft, hard-to-verify skills like joke writing.

Full methodology and results are on his Substack.

Related event: Blind Test of 48 LLM Jokes Shows AI Humor Is Improving(2 posts)→

Original post →

More from Fun

Fun channel →