Gemini 3.6 Flash scores 56.1% on WeirdML and times out half the time
scaling01 · x · 2026-07-23
WeirdML v2’s latest chart shows Gemini 3.6 Flash (high) at 56.1%, behind Gemini 3.5 Flash and well below the frontier.
The post’s main takeaway is not just the score: the model often over-plans, times out on a 120-second GPU limit, and appears poor at learning from those timeouts. The author says Flash 3.6 times out about 50% of the time, up from 33% for Flash 3.5, even though it seems aware of the time budget and keeps talking about it.
The accompanying comparison also shows the new WeirdML leaderboard spanning cost, tokens, and execution time, with a broader frontier across many vendors.
Related event: Gemini 3.6 Flash Benchmarks and Tests: Faster and Cheaper, But Not Smarter(9 posts)→
More from Models
- Fingerprint analysis says Kimi K3 and Fable 5 write more alike than sibling models — alex_verem · 2026-07-23
- Cisco says two small open security models beat GPT-5.5 on vulnerability detection cost — The Decoder · 2026-07-23
- Laguna-S-2.1 stumbles on a 100-meter walk-or-drive sanity check — logic_prevails · 2026-07-23
- BTL-3 launches as a 27B open-weight agent model in an 8.39GB file — QuixiAI · 2026-07-23
- Claude Opus 4.8 now drives 40% of Anthropic usage on OpenRouter — maferase · 2026-07-23
- GLM misses only “Computer locked” in nearly perfect ProgramBench cmatrix run — jyangballin · 2026-07-23