SemiAnalysis: Gemini 3.8 Flash and Muse Spark 1.3 look clearly benchmaxxed
firstadopter · x · 2026-09-08
- SemiAnalysis argues Gemini 3.8 Flash and Muse Spark 1.3 are among the most clearly benchmaxxed models yet: comparable to GPT-6 and Fable 5.1 on Terminal Bench 2.1, but markedly worse on Terminal Bench 4.0.
- firstadopter quote-shares the take, quipping that Google's benchmaxxing reputation is unmatched and asking where Gemini 3.5 Pro is.
More from Models
- Qwen3.5 35B-A3B sampling differs between Tinker API and Alibaba Cloud — philhchen · 2026-09-08
- GPT-6 Astra scores 95% on robot control task — but critics say demos are gamed — GaryMarcus · 2026-09-08
- Power user burns through Pro limits on Astra, verdict: not AGI, just a bigger code monkey — GaryMarcus · 2026-09-08
- Epoch: Astra beats most humans at card game but shows little continual learning — teortaxesTex · 2026-09-08
- GPT-6 Astra usage limits have been reset for paid users — xiaohu · 2026-09-08
- Astra's computer use is strikingly fast, likely planning steps in batch — shekitup · 2026-09-08