Dev calls 3D render evals 'mid': rerunning the same prompt beats any model gap
BLUECOW009 · x · 2026-09-23
Developer BLUECOW009 argues the 3D render evals circulating on X are among the most meaningless benchmarks: simply sampling the same prompt multiple times yields visibly better results, so scores reflect luck more than real capability differences between models.
More from Models
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23