Meta's Muse Spark 1.1 APEX-Agents Score Jumps to 41.9% After Filter Tuning
RylanSchaeffer · x · 2026-07-22
Mercor released an updated score for Meta's Muse Spark 1.1 on the APEX-Agents benchmark.
Previously, the model scored 0 on roughly 10% of long-horizon professional tasks (in banking, law, and consulting) due to false positives in content moderation. After collaborating with Meta to tune the safety filters on the public checkpoint, its Pass@1 score increased from 37.1% to 41.9%. This places the model at #2 overall, just behind Claude Fable 5 (43.3%).
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27