One agent loop, 14 models, hard caps: lessons from a photo-to-Blender benchmark

smith2008 · reddit · 2026-10-02

The author benchmarked 14 models as the brain of a photo-to-Blender agent under identical hard caps: 20 minutes, $4, and 60 requests per scene, with a deterministic judge scoring re-rendered scenes.

Key lessons:

Top of the board: GPT-6 Astra 66 ($3.91/attempt), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Claude Sonnet 5.5 56 ($0.50).

Related event: GPT-6 Astra Sweeps 14-Model 3D Reconstruction Benchmark(3 posts)→

Original post →

More from coding & agent

coding & agent channel →