MirrorCode Leaderboard Adds Gemini 3.1 Pro Evaluations
scaling01 · x · 2026-08-12
The MirrorCode coding leaderboard has been updated to include Gemini 3.1 Pro. The evaluation tasks are massive, featuring a budget of 10 billion tokens and a 7-day limit per attempt. Results for more models will be added as they complete.
More from Models
- Users Accuse Anthropic of Cooked Evals, Claiming Real API Performance Lags — GabGarrett · 2026-08-12
- Encrypted Chain-of-Thought in Proprietary LLMs Can Be Extracted via Weaker Sibling Models — Simon Willison · 2026-08-12
- GPT vs Claude: Which Model Better Understands Your Intent? — nickbaumann_ · 2026-08-12
- Expert Warns AI Watermarks Could Increase Code Entropy, Harming Coding Agents — wightmanr · 2026-08-12
- e/acc's Beff Jezos flags a 'vibe shift' as more users daily drive Grok — beffjezos · 2026-08-12
- Kimi K3 Distillation Controversy: Authors Admit No Proof, Likely Data Contamination — bookwormengr · 2026-08-12