Internal benchmark: GLM-5.3 Flash writes scripts at 1/429th the price, 1.7 points behind

OnlyProggingForFun · reddit · 2026-09-27

The author ran an internal creative-writing benchmark: models wrote full YouTube scripts (10 real tasks, 5 scripts each) scored by three AI judges against their own edited references. GLM-5.3 Flash cost $0.0074 per script and scored 88.2, vs Claude Fable 5.1 at max effort scoring 89.7 for $3.15 — 429x the price for 1.7 points. Many pricier setups scored worse. The author suggests blind-testing with hidden model names and counting publishable scripts and editing time.

Related event: Homegrown writing benchmark: GLM-5.3 Flash near top quality at tiny cost(2 posts)→

Original post →

More from Models

Models channel →