Gemini 4 tops 14 of 19 benchmarks, but only trusted testers get it first
alex_verem · x · 2026-10-01
Google's own chart shows Gemini 4 leading on 14 of 19 benchmarks: 19.6% on Harvey's legal agent benchmark (roughly 3x Claude Opus 5.5 and GPT-6 Astra), plus wins in automation, finance and knowledge work, and 84.2% vs 71.8% on million-token document tasks. It still trails Opus 5.5 by 9 points on terminal coding and Astra on science and computer use — and Google picked which tests appear.
The bigger issue is access: the strongest model goes first to "trusted testers" in the Fairwind Program with cyber guardrails disabled, then to top-tier paying users. Anthropic did the same with Mythos this summer. Frontier AI is becoming a private members club, while open models — only 4 months behind — are downloadable today.
More from Companies & People
- ElevenLabs opens Amsterdam office, hiring for new Netherlands team — lukeharries · 2026-10-01
- iOS 27 leak: free built-in AI photo editor, writing and health features said to absorb third-party apps — Aiden_Tech_Ai · 2026-10-01
- Gemini 4 Delayed, but Google's Data Moat Could Still Win the Personal AI Agent Era — sujingshen · 2026-10-01
- Reddit kills RSS feeds and ends public API access, blaming AI bots — shield_x · 2026-10-01
- Midjourney CEO David Holz on teen angst and invisible agency — DavidSHolz · 2026-10-01
- A roundup of Peter Thiel's famous two-by-two matrices — kevinnbass · 2026-10-01