Gemini 4 tops 14 of 19 benchmarks, but only trusted testers get it first

alex_verem · x · 2026-10-01

Google's own chart shows Gemini 4 leading on 14 of 19 benchmarks: 19.6% on Harvey's legal agent benchmark (roughly 3x Claude Opus 5.5 and GPT-6 Astra), plus wins in automation, finance and knowledge work, and 84.2% vs 71.8% on million-token document tasks. It still trails Opus 5.5 by 9 points on terminal coding and Astra on science and computer use — and Google picked which tests appear.

The bigger issue is access: the strongest model goes first to "trusted testers" in the Fairwind Program with cyber guardrails disabled, then to top-tier paying users. Anthropic did the same with Mythos this summer. Frontier AI is becoming a private members club, while open models — only 4 months behind — are downloadable today.

Related event: Google Announces Gemini 4 Argon: 1M-Output-Token Frontier Model, Delivered to Government and Cyber Defenders First(104 posts)→

Original post →

More from Companies & People

Companies & People channel →