Internal benchmark: GLM-5.3 Flash writes scripts at 1/429th the price, 1.7 points behind
OnlyProggingForFun · reddit · 2026-09-27
The author ran an internal creative-writing benchmark: models wrote full YouTube scripts (10 real tasks, 5 scripts each) scored by three AI judges against their own edited references. GLM-5.3 Flash cost $0.0074 per script and scored 88.2, vs Claude Fable 5.1 at max effort scoring 89.7 for $3.15 — 429x the price for 1.7 points. Many pricier setups scored worse. The author suggests blind-testing with hidden model names and counting publishable scripts and editing time.
Related event: Homegrown writing benchmark: GLM-5.3 Flash near top quality at tiny cost(2 posts)→
More from Models
- Google's AI lineup looks stalled: nano banana untouched since June, Gemini Pro since February — haider1 · 2026-09-27
- Rumor: partners got a better Claude Sonnet 5 checkpoint, release expected next week — kimmonismus · 2026-09-27
- Codex surprise hard reset angers users: 60% saved quota wiped, reset pushed 7 days out — ChrisUniverse · 2026-09-27
- Ex-NVIDIA engineer: US labs ignored global users, so the world runs on Chinese models — ivan_bezdomny · 2026-09-27
- GPT-OSS chat template bug silently drops past answers, degrading multi-turn coherence — arbv · 2026-09-27
- ScienceArena benchmark: LLMs score 64.5% on chemistry tasks needing structural diagrams vs 74.1% without — geoffwolfe · 2026-09-27