Evaluating 14 LLMs as visual coding agents: re-rendering, a deterministic judge, provider quirks
smith2008 · reddit · 2026-10-02
The author details the evaluation design behind a photo-to-Blender agent benchmark for 14 models:
- Re-render every delivered scene yourself; don't score what the model hands in.
- No LLM judge: silhouette overlap (30%), edge F1 (40%), multiscale RGB error (15%), mesh health (15%), mapped to 0-100 via calibrated scales.
- Anti-cheating: a second 35° camera swing exposes flat stand-ins; referenced images are digest-compared against the reference.
- Failures stay in the average; zero-score runs are included.
- Equal plumbing: cache-control for Anthropic, a closing user turn for Google — before fixes, Claude ran 0% cached vs 96% for OpenAI.
Results: GPT-6 Astra 66 ($3.91), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Sonnet 5.5 56 ($0.50), GLM 5.3 Flash 36 ($0.089); DeepSeek V4.1 Flash and Qwen3.8 Max scored 0 (nothing saved in 20 minutes).
Next steps: three runs per image for the top group and a model-neutral agent loop.
More from coding & agent
- Mirelo brings AI sound effects generation to AWS's Kiro agentic IDE — ordax · 2026-10-02
- Using Codex 8 hours a day and still can't burn through the limits — honkballs · 2026-10-02
- Free 18k-star GitHub course on Harness Engineering: 14 lectures + 8 hands-on projects — Hesamation · 2026-10-02
- Manifesto: open-source project makes UI and agents share the same app-owned actions — TraditionalListen994 · 2026-10-02
- Debian kernel alert teems with void* bugs: who (or what) is auditing the Linux kernel? — mircomusolesi · 2026-10-02
- Google's A2A protocol sees hype but little production use yet — Diarnstinc_Crew_6510 · 2026-10-02