Opus 5 adds numeric self-checks to the Boeing benchmark and outbuilds Fable
victormustar · x · 2026-07-25
- The post argues that a comparison video makes it easier to see that Opus 5 produced a more detailed result than Fable on the Boeing benchmark.
- The quoted benchmark notes that Fable 5 took 42 minutes and produced about 4 files / 1,000 lines, using a Puppeteer harness, 12 camera presets, 17 render passes, and 54 screenshot inspections.
- Opus 5 took 71 minutes, but expanded the work to 20 modules / 4,000 lines and used a second tool, measure.mjs, to extract quantitative geometry from the running scene.
- That tool measured 14 aircraft metrics — including span, wheelbase, wing root leading-edge station, sweep, dihedral, and gear track — and compared them against published 747-400 figures with per-metric tolerances.
- The image and commentary together suggest Opus 5 does not just visually verify the scene; it tries to measure and self-check the generated output numerically.
Related event: Opus 5 Shows Superior Detail in Boeing Render Comparison(2 posts)→
More from Models
- Miles Brundage says Anthropic’s chart issue is really about over-trusting Claude — Miles_Brundage · 2026-07-25
- A benchmark chart becomes an AI meme after viewers spot the messy numbers — Miles_Brundage · 2026-07-25
- FrontierCode 1.1 shows Opus 5 can score lower under stricter reasoning settings — andrew_n_carr · 2026-07-25
- Bug Hunt Bench: GPT-5.6 Sol fixes 22 bugs, Opus 5 12, on a 45-bug repo — PawelHuryn · 2026-07-25
- Claude Opus 5 builds a Rocket League clone on just 27% of a Max plan — soumitrashukla9 · 2026-07-25
- Repost accuses Claude benchmark chart of highlighting only GPT-5.6 Sol’s sole win — soumitrashukla9 · 2026-07-25