Meta's Muse Spark 1.2 beats GPT-5.5 on math benchmark, slightly behind Kimi K3
alexandr_wang · x · 2026-08-08
According to ErdosBench evaluation, Meta's new model Muse Spark 1.2 solved 40 out of 226 research-level math problems, outperforming GPT-5.5 xhigh and slightly behind Kimi K3. The model shows good proof hygiene, high B-grade review yield, no rejected strong claims, but fewer decisive A-grade closures.
More from Models
- Why LLMs Can't Count Tokens: The Need for Intermediate Steps — ctjlewis · 2026-08-08
- DeepSeek V4 Flash Dominates ARC-AGI Cost-Performance Frontier at 1/4 the Cost — GregKamradt · 2026-08-08
- Databricks Reveals Enterprise AI Coding Economics: Newer Models Aren't Always Cheaper — Yuchenj_UW · 2026-08-08
- Matt Shumer's Tips for Claude Opus 5: Clear Presets and Let Go of Control — mattshumer_ · 2026-08-08
- DeepSeek V4 Flash launches on engy.ai, 68% cheaper than official API — markjeffrey · 2026-08-08
- Are MoEs Completely Overrated? Reddit User Slams Low Execution Intelligence — infieldmitt · 2026-08-08