Gemini 3.6 Flash scores 56.1% on WeirdML and times out half the time

scaling01 · x · 2026-07-23

WeirdML v2’s latest chart shows Gemini 3.6 Flash (high) at 56.1%, behind Gemini 3.5 Flash and well below the frontier.

The post’s main takeaway is not just the score: the model often over-plans, times out on a 120-second GPU limit, and appears poor at learning from those timeouts. The author says Flash 3.6 times out about 50% of the time, up from 33% for Flash 3.5, even though it seems aware of the time budget and keeps talking about it.

The accompanying comparison also shows the new WeirdML leaderboard spanning cost, tokens, and execution time, with a broader frontier across many vendors.

Related event: Gemini 3.6 Flash Benchmarks and Tests: Faster and Cheaper, But Not Smarter(9 posts)→

Original post →

More from Models

Models channel →