Opus 5.5 reportedly beats professional human baseline on full-screenplay writing benchmark
141_1337 · reddit · 2026-09-25
A Reddit user reports that Opus 5.5 surpassed the professional human baseline on a benchmark testing whether AI can write complete scripts — judged on voice, substance, pacing, hooks, and minimal "AI slop." Screenshots are attached in the post, though the benchmark's methodology and provenance remain unverified; treat as a capability observation.
More from Models
- Dev Compares Codex vs Claude: Claude Nails Multi-Agent Orchestration, Codex Doesn't — madhavsinghal_ · 2026-09-25
- Student Pays $30/Month for Google AI Pro, Still Finds Gemini Unreliable and Hallucinatory — makeitbumthem · 2026-09-25
- Meta's Muse app hit 1.8M iOS downloads in 12 days, beating ChatGPT's 1.3M — Beth_Kindig · 2026-09-25
- Opus 5.5's token usage is far more reasonable: 13-hour session on a 20x Max plan — majidmanzarpour · 2026-09-25
- TypeSafe's Jev: A Model That Can't Write but Decides, Claiming 200x Speed and 400x Cost Savings — rseroter · 2026-09-25
- Prime Intellect pitches highest-throughput GLM 5.3 inference with OpenAI-compatible eval API — willcb · 2026-09-25