Gemini 3.8 Flash (high) hits 84.8% on WeirdML v2, first Flash to beat Gemini 3.1 Pro

teortaxesTex · x · 2026-10-01

htihle reports Gemini 3.8 Flash (high) scoring 84.8% on WeirdML v2 — on par with GPT 5.5 (xhigh) at a fraction of the cost, and the first Flash version to beat Gemini 3.1 Pro (72.1%).

The main gain is better feedback handling: it no longer insists on bloated pipelines that repeatedly time out, the key failure mode of earlier Flash versions. On WeirdML v3, Gemini clusters with K3 (slightly above) and below V4.1, which the author estimates would score 79-81% on v2. This is likely among the last v2 results published, with focus shifting to v2/v3 cross-comparison data.

Original post →

More from Models

Models channel →