Meta's Muse Spark 1.3 tops Gemini 3.8 Flash on most overlapping benchmarks, crushes long-context MRCR
ChrisGPT · x · 2026-09-03
Blogger ChrisGPT compiled benchmark comparisons of Meta's newly released Muse Spark 1.3 against Google's same-day Gemini 3.8 Flash (figures attributed to Google DeepMind, independently unverified).
- Muse Spark 1.3 wins most overlapping benchmarks: GDPVal-AA v2 1754 vs 1545, DeepSWE 75.4 vs 73.7, OSWorld 2.0 66.9 vs 59.0; Gemini edges Terminal-Bench 2.1 (89.4 vs 88.8).
- Long context is the standout: 98.5 on MRCR 256K–512K and 98.1 on 512K–1M, versus 91.5 and 73.8 for GPT-5.6 Sol.
- The author notes it beats Opus 5 on DeepSWE and Terminal-Bench, though Opus remains stronger on GDPVal and OSWorld.
More from Models
- GLM 5.3 and 5.3 Flash Now Free to Try on Together Chat, No API Setup — oilmutt · 2026-09-03
- LatchBio finds Grok's refusals come from the model itself, while rivals rely on external safety layers — kenbwork · 2026-09-03
- Gemini 3.8 Flash Accused of Bench Overfitting, Regressing vs 3.7 in Third-Party Tests — bindureddy · 2026-09-03
- Gemini 3.8 Flash shows double-digit lift in user satisfaction over 3.7 — tokumin · 2026-09-03
- Muse Spark 1.3 debuts at #3, first model to slot between Claude and GPT — alexandr_wang · 2026-09-03
- Google Researcher Mocks Astra Thinking-Token Outcry as Manufactured Angst — rao2z · 2026-09-03