GPT-5.6 Additional Visual Benchmark Results
Afinetheorem · x · 2026-07-10
The author added that Terra is roughly on par with Gemini 2.0 Flash, and Luna is around the level of Sonnet 3.5/Grok 4.5. This indicates subpar performance on this specific benchmark, making it one of the worst OpenAI models the author has tested.
Related event: EyeBench-v3 Updates: Sol Model Takes the Lead(4 posts)→
More from Models
- Gemini 3.6 Flash goes live in Antigravity with 17% fewer output tokens — rseroter · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Gemini 3.6 Flash benchmark results reignite concerns that Google is slipping behind — minxio_ · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22