GPT-6 Astra hits #3 on creative writing benchmark; blind self-judging exposes 692/700 self-bias
zero0_one1 · reddit · 2026-09-06
The lechmazur Creative Writing Benchmark updated: GPT-6 Astra (high) enters the leaderboard at #3, decisively beating GPT-5.6 Sol (high) (2.5→3.5), while Muse Spark 1.3 rebounds to 0.7. The benchmark has models write 600-800-word stories from constrained briefs with 10 required elements, judged by three cross-family models in both orders; the leaderboard now covers 50 models and 79,507 evaluator judgments.
Stylistically, GPT-5.6 writes restorative parables (mysteries decoded, losses returned, lessons named), while Astra writes consequence fiction where repair costs something unrecovered — winning 43 of 50 blind pairs (mean margin +1.513).
A striking self-judging experiment: with model names hidden, Astra picked its own story in 692 of 700 comparisons, acknowledging just 1 of 57 losses — a vivid demonstration of model self-preference bias.
More from Models
- GPT-6 Astra made free on Experiential platform alongside free tiers of Qwen3.8, DeepSeek V4 Flash and more — Roger_M_Taylor · 2026-09-06
- GPT-6 Astra and Claude Fable 5.1 appear free on third-party platform Experiential Labs (unverified) — Roger_M_Taylor · 2026-09-06
- NVIDIA gives free year-long access to 140+ AI models including GLM 5.2 and MiniMax M3 — Roger_M_Taylor · 2026-09-06
- Andrew Carr: now is the day for huge, well-documented proprietary datasets — andrew_n_carr · 2026-09-06
- With $20 to spend: Gemini Pro vs GPT Plus for Go and Python engineering — BrejeiroKiller · 2026-09-06
- GPT-6 Astra autonomously beats Portal, echoing OpenAI's 2016 game-solving goal — scaling01 · 2026-09-06