Gemini 3.8 Flash (high) hits 84.8% on WeirdML v2, first Flash to beat Gemini 3.1 Pro
teortaxesTex · x · 2026-10-01
htihle reports Gemini 3.8 Flash (high) scoring 84.8% on WeirdML v2 — on par with GPT 5.5 (xhigh) at a fraction of the cost, and the first Flash version to beat Gemini 3.1 Pro (72.1%).
The main gain is better feedback handling: it no longer insists on bloated pipelines that repeatedly time out, the key failure mode of earlier Flash versions. On WeirdML v3, Gemini clusters with K3 (slightly above) and below V4.1, which the author estimates would score 79-81% on v2. This is likely among the last v2 results published, with focus shifting to v2/v3 cross-comparison data.
More from Models
- OpenAI says Moonshot AI-linked individuals led campaign making 16,000 attempts to extract hidden reasoning — kimmonismus · 2026-10-01
- True Positive Weekly #180: Xiaomi's MIT-licensed MiMo-V2.6, physicist-style LLM pruning, OpenHands — burkov · 2026-10-01
- Perplexity's pplx-embed-v2-context-9b-preview sets SOTA on ConTEB benchmark — perplexity_ai · 2026-10-01
- OpenAI Accuses Chinese Lab of Distillation: 16,000 Requests in 2 Days to Extract Hidden Reasoning — rohanpaul_ai · 2026-10-01
- AI Hype Cycles: The Same Crowd Bashing OpenAI Just Spent Months Bashing Anthropic — haider1 · 2026-10-01
- Player trusts Gemini on its own TTS pricing, gets billed 70p after Gemini said it would be free — Environmental_Ad3162 · 2026-10-01