Meta's Muse Spark 1.3 debuts third on Stata benchmark, behind only Claude
alexandr_wang · x · 2026-09-03
Early eval results from the Stata benchmark shared by Alexandr Wang show Meta's Muse Spark 1.3 debuting as the third top-performing model, beating every model except the top-ranked Claude Fable. Details and the full leaderboard have not yet been published, so treat with caution.
More from Models
- How Could Chinese Open-Source Models Actually Hurt the US? Two Failure Modes Explained — matanSF · 2026-09-03
- Meta's Muse Spark 1.3 tops Gemini 3.8 Flash on most overlapping benchmarks, crushes long-context MRCR — ChrisGPT · 2026-09-03
- Insider Praises Gemini 3.8 Flash: Better at Requirements, More Critical, Fast — prajdabre · 2026-09-03
- Tested GLM-5.3 abliterated model: 4x the cost, worse performance than the original — BLUECOW009 · 2026-09-03
- Gemini 3.8 Flash Accused of Benchmark Overfitting, Regressing vs 3.7 in Independent Tests — bindureddy · 2026-09-03
- Submission timelines hint at long-horizon post-training: open models submit >10 hours late — nrehiew_ · 2026-09-03