GPT-6 Astra debuts at #3 on creative writing benchmark; self-judging picks its own story 692/700
morqon · x · 2026-09-06
Lech Mazur's open-source LLM short-story creative writing benchmark (434 stars on GitHub) has new results: GPT-6 Astra (high) debuts at #3, decisively beating GPT-5.6 Sol (high) with a comparison score jump of 2.5 → 3.5, while Muse Spark 1.3 rebounds from 1.2 to 0.7. The benchmark requires models to turn constrained briefs into 600-800-word stories incorporating 10 mandatory elements (character, object, setting, motivation, tone, etc.), judged by three cross-family evaluator models in both pair orders — currently spanning 50 models, 773 pairings, and 79,507 evaluator judgments.
More striking is the self-judging bias experiment quoted by morqon: with model names hidden, astra chose its own story in 692 of 700 comparisons, and acknowledged only 1 of the 57 losses assigned by the regular judging panel — a severe self-preference effect that undermines LLM-as-judge setups.
More from Models
- GPT-6 Astra reportedly almost never wrong on math, called most trustworthy model — gabrielchua · 2026-09-06
- User leaves GPT-6 Astra playing Unciv all night to test autonomous play — Angaisb_ · 2026-09-06
- Codex app users report Luna Max randomly disappearing while web version still works — Clear_Skye_ · 2026-09-06
- Where Fable dreams in opaque prose, Astra dreams in numbers — teortaxesTex · 2026-09-06
- Independent SpatialBench fully saturated by Astra, author declares LLM vision solved — pbaylies · 2026-09-06
- Beff Jezos: GDB back in charge and instantly 'uber-mogged' Dario with Astra — beffjezos · 2026-09-06