Experiment finds newer LLMs are worse at having fun

An experiment by David Holz let multiple LLMs play freely and rate each other's fun, finding counterintuitively that newer models play less well, with Claude Opus 4.6 repeatedly winning.

2026-10-05 ~ 2026-10-05 · 4 related posts