Gemini 3.6 Flash scores 56.1% on WeirdML and times out half the time
scaling01 · x · 2026-07-23
WeirdML v2’s latest chart shows Gemini 3.6 Flash (high) at 56.1%, behind Gemini 3.5 Flash and well below the frontier.
The post’s main takeaway is not just the score: the model often over-plans, times out on a 120-second GPU limit, and appears poor at learning from those timeouts. The author says Flash 3.6 times out about 50% of the time, up from 33% for Flash 3.5, even though it seems aware of the time budget and keeps talking about it.
The accompanying comparison also shows the new WeirdML leaderboard spanning cost, tokens, and execution time, with a broader frontier across many vendors.
Related event: Gemini 3.6 Flash Review: Faster and Cheaper, But Not Smarter(16 posts)→
More from Models
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11