Blind test: Opus 5.5 vs Sonnet 5.5 each film a 70s-style industrial explainer from the same brief
tobowers · x · 2026-10-08
Developer Topper Bowers had Opus 5.5 and Sonnet 5.5 each produce a short film from the same brief — a 70s/80s industrial explainer featuring generated video (H3 on fal), HyperFrames and Creative Loom — using the same repo, skills and Loom tools. Sonnet rebuilt it as a blind subagent, barred from Opus's folders.
- Output: Opus made 103s / 9 shots in 25 min plus one human feedback round; Sonnet made 82s / 8 shots in 24 min with zero feedback — every image, clip and line accepted first take
- Cost: $0.30 total; Sonnet's 8 clips ≈ $0.22
- Price discrepancy: Opus re-checked fal's live page mid-run and used $0.015/s; Sonnet relied on the repo's stale fal table from Sep 15 ($0.025/s)
- Includes frame-strip comparisons, a per-item brief checklist and full run narratives
Related event: Side-by-Side Test Shows Opus 5.5 Beats Sonnet 5.5 in Creative Tasks(2 posts)→
More from coding & agent
- Grok bot surfaces $25,730 in live GitHub bounties, top one pays $10k — prasenx · 2026-10-08
- Talk: how to RL-train an agent running inside a harness you didn't write — SergioPaniego · 2026-10-08
- 8 open-source AI agent tools to know, from computer-use to browser automation — nikola_mr64990 · 2026-10-08
- Test quality follows module design: test larger units instead of banning AI tests — mattpocockuk · 2026-10-08
- Matt Pocock: AI test quality depends on codebase design, banning tests is wrong — mattpocockuk · 2026-10-08
- Codex Spent 48 Hours Migrating WebGL to WebGPU by Patching Upstream Libraries — StewartalsopIII · 2026-10-08