Meta's Muse Spark 1.2 beats GPT-5.5 on math benchmark, slightly behind Kimi K3

alexandr_wang · x · 2026-08-08

According to ErdosBench evaluation, Meta's new model Muse Spark 1.2 solved 40 out of 226 research-level math problems, outperforming GPT-5.5 xhigh and slightly behind Kimi K3. The model shows good proof hygiene, high B-grade review yield, no rejected strong claims, but fewer decisive A-grade closures.

Original post →

More from Models

Models channel →