Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug

eyishazyer · x · 2026-09-11

After testing Tencent's Hy4 preview (which needed about an hour to build one correct Echo Maze game), the author ran the identical prompt through Claude Fable 5.1, GPT-5.6 Luna, and Gemini 3.6 Flash:

Expecting to just confirm Luna's bug, the author inspected all three codebases and found the same bug in Fable's code too. Gemini's walls were there the whole time — two colors just 17 points apart on a 255 scale, invisible on screen.

Takeaway: all three models finished in minutes what took a human an hour, yet each still required manual code review. Checking their homework remains mandatory.

Original post →

More from coding & agent

coding & agent channel →