Fable 5.1 Tops Adversarial LLM Debate Benchmark as GLM-5.3 Debuts at #4

teortaxesTex · x · 2026-09-06

LechMazur updated his LLM Debate Benchmark, which tests how well models sustain arguments under adversarial, multi-turn opposition across hundreds of topics. Each matchup runs twice with sides swapped, judged by a three-model panel, and ranked via Bradley-Terry ratings centered near 1500.

Key results:

Poster teortaxesTex frames this as a style split between "vicious wordcel" Fable and the "innocent child genius" Astra. The benchmark is open-sourced on GitHub (lechmazur/debate).

Related event: Claude Fable 5.1 Tops LLM Debate Benchmark, GLM-5.3 Debuts in Top Four(2 posts)→

Original post →

More from Models

Models channel →