Muse Spark 1.3 shines on brand-new, unoptimizable CursorBench 4.0
jyangballin · x · 2026-09-11
- leerob announced CursorBench 4.0, adding new tasks on instruction following and long-horizon work on challenging projects, and raising difficulty so all models score lower — a benchmark models couldn't have optimized against.
- mattdeitke argued Muse Spark 1.3 performing strongly on this fresh benchmark shows real generalization, countering Meta's old "benchmaxxing" reputation.
More from Models
- OpenAI pauses new $200 ChatGPT Pro sign-ups as heavy users burn through quota — MickeySteamboat · 2026-09-11
- Genspark launches Gen-1 Slides, a work model priced at 1/17th of Opus 5 — Scobleizer · 2026-09-11
- OpenAI's Rumored Internal Model 'Bel' Could Be the Most Hyped Release Ever — imadade · 2026-09-11
- DeepSeek-V4.1-Flash rolls out to Ollama cloud Pro subscribers after Max and Team debut — ollama · 2026-09-11
- GPT-6 Astra called a computer-use model built for knowledge work like accounting — mckbrando · 2026-09-11
- Sakana AI launches Fugu Max and Fugu Ultra v2, matching elite models at 2-6x lower cost — SakanaAILabs · 2026-09-11