SemiAnalysis: Gemini 3.8 Flash and Muse Spark 1.3 are among the most benchmaxxed models yet
Neurogence · reddit · 2026-09-09
SemiAnalysis states that Gemini 3.8 Flash and Muse Spark 1.3 are 'two of the most clearly benchmaxxed models we've seen yet,' implying a notable gap between benchmark scores and real-world capability. The post shares the claim with an accompanying chart image.
Related event: SemiAnalysis Calls Gemini 3.8 Flash a Benchmark-Gaming Model(3 posts)→
More from Models
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11
- Benchmark author says OpenRouter unreliably honors Meta Muse effort levels, EU payments broken — PawelHuryn · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11