Gemini 3.8 Flash aces benchmarks but can't tell when a simple query needs web search
Distinct_Fox_6358 · reddit · 2026-09-04
A Reddit user points out that Gemini 3.8 Flash, despite scoring extremely highly on benchmarks, fails to recognize when even a simple question requires web search — highlighting the growing gap between leaderboard scores and real-world tool-use behavior.
More from Models
- DeepSeek V4 Flash Vision impresses via API, but 305B needs 4x GB300 to run — No_Issue_8224 · 2026-09-04
- Don't fall for GPT-6's 98.6% ARC AGI-3: Nvidia AVO already hit 100% — Informal-Trouble2183 · 2026-09-04
- Multiverse Computing launches Quasar 438B, strongest European model at 438B params — teortaxesTex · 2026-09-04
- Skeptical take: OpenAI can't train large models, pivots to RL and inference — teortaxesTex · 2026-09-04
- mark_k: pointless to start new coding projects with lesser AIs before Astra access — mark_k · 2026-09-04
- GPT-6 Astra reviews: half the per-task cost, but CoT monitoring breaks down — vista8 · 2026-09-04