GPT-6 Astra debuts at #3 on creative writing benchmark; self-judging picks its own story 692/700

morqon · x · 2026-09-06

Lech Mazur's open-source LLM short-story creative writing benchmark (434 stars on GitHub) has new results: GPT-6 Astra (high) debuts at #3, decisively beating GPT-5.6 Sol (high) with a comparison score jump of 2.5 → 3.5, while Muse Spark 1.3 rebounds from 1.2 to 0.7. The benchmark requires models to turn constrained briefs into 600-800-word stories incorporating 10 mandatory elements (character, object, setting, motivation, tone, etc.), judged by three cross-family evaluator models in both pair orders — currently spanning 50 models, 773 pairings, and 79,507 evaluator judgments.

More striking is the self-judging bias experiment quoted by morqon: with model names hidden, astra chose its own story in 692 of 700 comparisons, and acknowledged only 1 of the 57 losses assigned by the regular judging panel — a severe self-preference effect that undermines LLM-as-judge setups.

Related event: GPT-6 Astra Debuts at No.3 on Creative Writing Benchmark, Blind Tests Favor Itself(2 posts)→

Original post →

More from Models

Models channel →