One agent loop, 14 models, hard caps: lessons from a photo-to-Blender benchmark
smith2008 · reddit · 2026-10-02
The author benchmarked 14 models as the brain of a photo-to-Blender agent under identical hard caps: 20 minutes, $4, and 60 requests per scene, with a deterministic judge scoring re-rendered scenes.
Key lessons:
- Enforce limits outside the agent and require early saves — runs ending with nothing saved score 0.
- Models disagree about being done: GPT-5.6 Sol/Terra quit after 8 minutes with most budget unspent; GPT-6 Astra refined until the $4 cap every time and won all three photos.
- Heavy reasoners lose to the clock: DeepSeek V4.1 Flash and Qwen3.8 Max spent 83%/68% of output reasoning and saved nothing in 20 minutes.
- Provider plumbing matters: Anthropic needed explicit cache markers (0% → 96-97% cached), a Google endpoint needed a closing user turn, and Kimi K3 once returned an empty first answer.
Top of the board: GPT-6 Astra 66 ($3.91/attempt), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Claude Sonnet 5.5 56 ($0.50).
Related event: GPT-6 Astra Sweeps 14-Model 3D Reconstruction Benchmark(3 posts)→
More from coding & agent
- Mirelo brings AI sound effects generation to AWS's Kiro agentic IDE — ordax · 2026-10-02
- Using Codex 8 hours a day and still can't burn through the limits — honkballs · 2026-10-02
- Free 18k-star GitHub course on Harness Engineering: 14 lectures + 8 hands-on projects — Hesamation · 2026-10-02
- Manifesto: open-source project makes UI and agents share the same app-owned actions — TraditionalListen994 · 2026-10-02
- Debian kernel alert teems with void* bugs: who (or what) is auditing the Linux kernel? — mircomusolesi · 2026-10-02
- Google's A2A protocol sees hype but little production use yet — Diarnstinc_Crew_6510 · 2026-10-02