Opus 5.5 reportedly beats professional human baseline on full-screenplay writing benchmark

141_1337 · reddit · 2026-09-25

A Reddit user reports that Opus 5.5 surpassed the professional human baseline on a benchmark testing whether AI can write complete scripts — judged on voice, substance, pacing, hooks, and minimal "AI slop." Screenshots are attached in the post, though the benchmark's methodology and provenance remain unverified; treat as a capability observation.

Original post →

More from Models

Models channel →