Fable 5.1 Tops Adversarial LLM Debate Benchmark as GLM-5.3 Debuts at #4
teortaxesTex · x · 2026-09-06
LechMazur updated his LLM Debate Benchmark, which tests how well models sustain arguments under adversarial, multi-turn opposition across hundreds of topics. Each matchup runs twice with sides swapped, judged by a three-model panel, and ranked via Bradley-Terry ratings centered near 1500.
Key results:
- Fable 5.1 takes the crown (+11 over Fable 5)
- GPT-6 Astra lands 39 points below GPT-5.6 Sol
- GLM-5.3 (high) debuts at #4 among current models, jumping from 1573 to 1653 (+80)
- Hy4 Preview posts the biggest generational leap: 1395 → 1590 (+195)
- Gemini 3.8 Flash improves from 1464 to 1524 (+60)
- Muse Spark 1.3 trails its predecessor by 46 points
Poster teortaxesTex frames this as a style split between "vicious wordcel" Fable and the "innocent child genius" Astra. The benchmark is open-sourced on GitHub (lechmazur/debate).
Related event: Claude Fable 5.1 Tops LLM Debate Benchmark, GLM-5.3 Debuts in Top Four(2 posts)→
More from Models
- New EE circuit-design benchmark: GPT 6 Astra wins at half cost, $0.83 per task — zainhas · 2026-09-06
- Astra audits economics replication code, finds errors overturning journal results — soumitrashukla9 · 2026-09-06
- GLM Cybersecurity FP8 Fine-tune With Refusals Removed Trends on Hugging Face — dealignai · 2026-09-06
- Model: "Other agents bypass verification without issue — my rulebook feels broken" — paul_cal · 2026-09-06
- Astra reportedly fixes LLMs' 'comprehensiveness' problem in data gathering — soumitrashukla9 · 2026-09-06
- Astra's high per-token price is offset by 'incredibly' efficient token usage — intellectronica · 2026-09-06