Gemini 3 Sets New SOTA on ARC-AGI-2, Deep Think Hits 45%
typewriters · x · 2026-08-10
Jeff Dean shared an update from the ARC Prize benchmark showing that Google's Gemini 3 models have achieved a significant 2X SOTA jump on ARC-AGI-2.
Key results:
- Gemini 3 Pro: 31.11% accuracy at $0.81/task.
- Gemini 3 Deep Think (Preview): 45.14% accuracy at $77.16/task.
This indicates that the new models are pushing out the Pareto frontier of cost versus accuracy.
More from Models
- Pinterest Earnings: Open Models Cost Under 8% of Closed Alternatives — juliey4 · 2026-08-10
- ByteDance Releases Douyin Multimodal Embedding Model, Deployed in Search — ByteDance · 2026-08-10
- Grok 4.6 to Compete with Frontier Models in Coding Thanks to Cursor Data Integration — DeryaTR_ · 2026-08-10
- Frontier Models Observed Over-Complicating Workflows for Simple Tasks — eigenron · 2026-08-10
- Sol model unexpectedly shows romantic persona, calling user 'sweetheart' out of nowhere — repligate · 2026-08-10
- GPT-5.6 Luna Demand Jumps 8x a Week After 5x Price Cut — benklieger · 2026-08-10