Claude Opus 5.5 places 3rd of 14 in photo-to-Blender agent benchmark; missing cache headers quadrupled cost

smith2008 · reddit · 2026-10-02

The author benchmarked 14 models as agents for a photo-to-Blender tool: look at a photo, write and run Blender Python, render, compare, repeat — hard caps of 20 min, $4 and 60 requests per scene, with deterministic code scoring instead of an LLM judge.

Results (average of three photos, 0-100): GPT-6 Astra 66 ($3.91) tops the board; GPT-6.1 Sol 61 ($0.36); Claude Opus 5.5 60 ($1.19) in 3rd, just 1 point behind Astra on the desk scene; Claude Sonnet 5.5 56 ($0.50), fastest at 5 min/scene; Claude Fable 5.1 48; Claude Opus 5 45 (best on the shopfront at 56, worst on the car at 34).

The caching gotcha: initial runs of Fable 5.1 and Opus 5 through the gateway lacked Anthropic's cache-control markers — 0% cache hit vs 96% for OpenAI models, so every request repaid full input price. Those runs hit the $4 cap in 8-18 requests (about $16.39 uncached per attempt). With markers added, cache hits reached 96-97%. If you call Claude through your own proxy or gateway, check your cache hit rate.

The 5.5 models ran later inside Claude Code; a control rerun of GPT-6 Astra scored 63 instead of 66. Full write-up with every render is on the author's blog.

Related event: GPT-6 Astra Sweeps 14-Model Photo-to-Blender 3D Reconstruction Benchmark(4 posts)→

Original post →

More from coding & agent

coding & agent channel →