Evaluating 14 LLMs as visual coding agents: re-rendering, a deterministic judge, provider quirks

smith2008 · reddit · 2026-10-02

The author details the evaluation design behind a photo-to-Blender agent benchmark for 14 models:

Results: GPT-6 Astra 66 ($3.91), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Sonnet 5.5 56 ($0.50), GLM 5.3 Flash 36 ($0.089); DeepSeek V4.1 Flash and Qwen3.8 Max scored 0 (nothing saved in 20 minutes).

Next steps: three runs per image for the top group and a model-neutral agent loop.

Original post →

More from coding & agent

coding & agent channel →