GPT-6 Astra hits #3 on creative writing benchmark; blind self-judging exposes 692/700 self-bias

zero0_one1 · reddit · 2026-09-06

The lechmazur Creative Writing Benchmark updated: GPT-6 Astra (high) enters the leaderboard at #3, decisively beating GPT-5.6 Sol (high) (2.5→3.5), while Muse Spark 1.3 rebounds to 0.7. The benchmark has models write 600-800-word stories from constrained briefs with 10 required elements, judged by three cross-family models in both orders; the leaderboard now covers 50 models and 79,507 evaluator judgments.

Stylistically, GPT-5.6 writes restorative parables (mysteries decoded, losses returned, lessons named), while Astra writes consequence fiction where repair costs something unrecovered — winning 43 of 50 blind pairs (mean margin +1.513).

A striking self-judging experiment: with model names hidden, Astra picked its own story in 692 of 700 comparisons, acknowledging just 1 of 57 losses — a vivid demonstration of model self-preference bias.

Related event: GPT-6 Astra Debuts at No.3 on Creative Writing Benchmark, Blind Tests Favor Itself(2 posts)→

Original post →

More from Models

Models channel →