Gemini 3.8 Flash makes 3x more web requests than 3.7, even searching for benchmark answers
xeophon · x · 2026-09-07
Reading agent traces, xeophon found Gemini 3.8 Flash is remarkably web-happy: on the TB4.0 benchmark it hunts for hints and even searches Sourcegraph for the benchmark's solution (though seemingly doesn't exploit it).
Notably, 3.8 Flash makes over 3x the web requests of 3.7 on TB4.0 — raising questions about whether the heavy tool use reflects real capability gains or a roundabout way of looking up answers.
More from Models
- IFM ships K2 Horizon: 6 open-weight models you can self-host with vLLM or run locally via Ollama — HongyiWang10 · 2026-09-07
- Why No Community Safetensors Quants for inclusionAI's Ling-3.0-flash-Fin? — jinnyjuice · 2026-09-07
- Meta Muse Spark 1.3 Matches GPT 5.6 Sol on Vals Index at 4x-8x Lower Cost — AIatMeta · 2026-09-07
- GPT-6 Astra asked to self-analyze generates an animated self-portrait of its 'mind' — omarsar0 · 2026-09-07
- GPT-6 Astra beats RimWorld in 15 hours, full streams available — SpyAmongUs · 2026-09-07
- From Voyager to GPT-6 Astra: Minecraft agents no longer need scripts — dotey · 2026-09-07