Meta's Muse Spark 1.1 APEX-Agents Score Jumps to 41.9% After Filter Tuning

RylanSchaeffer · x · 2026-07-22

Mercor released an updated score for Meta's Muse Spark 1.1 on the APEX-Agents benchmark.

Previously, the model scored 0 on roughly 10% of long-horizon professional tasks (in banking, law, and consulting) due to false positives in content moderation. After collaborating with Meta to tune the safety filters on the public checkpoint, its Pass@1 score increased from 37.1% to 41.9%. This places the model at #2 overall, just behind Claude Fable 5 (43.3%).

Original post →

More from Models

Models channel →