GPT-5.6 Additional Visual Benchmark Results

Afinetheorem · x · 2026-07-10

The author added that Terra is roughly on par with Gemini 2.0 Flash, and Luna is around the level of Sonnet 3.5/Grok 4.5. This indicates subpar performance on this specific benchmark, making it one of the worst OpenAI models the author has tested.

Related event: EyeBench-v3 Updates: Sol Model Takes the Lead(4 posts)→

Original post →

More from Models

Models channel →