Study: post-training makes LLMs funnier but homogenizes their jokes across 11 stages

Gold-Bat-3225 · reddit · 2026-09-30

A study tested whether post-training actually makes LLMs funnier, using open models with fully published training stages — Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B, and Qwen2.5 — tracking 11 stages with 100 joke prompts, 64 human raters, and 2,330 head-to-head judgments.

Key findings:

Method: humans judged base vs. final models; a calibrated model judge compared intermediate stages. Full report: laugh.so/research/humor-tax.

Original post →

More from Research

Research channel →