"Do better!" prompting stalls fast; even top VLMs understand images unevenly

keenanisalive · x · 2026-10-10

The author found that generic "do better!" prompts stalled quickly, so he switched to giving human feedback instead. His experience: even the best VLMs understand images unevenly at best, especially for high-end 3D and image artifacts—making visual-feedback loops a bottleneck for refining LLM-generated 3D models.

Original post →

More from coding & agent

coding & agent channel →