Muse Spark 1.1 Shows Significant Coding Improvements
EdwardSun0909 · x · 2026-07-11
Summarized evaluations indicate that Muse Spark 1.1 improved its Intelligence Index by 8 points over 1.0, primarily in scientific reasoning, coding, and knowledge.
Specific changes include:
- Coding Index: 59 → 71 (+12)
- SciCode: 52% → 58% (+6)
- Humanity's Last Exam: 40% → 45% (+5)
- AA-Omniscience: 4 → 18 (+14)
- GDPval-AA v2: 1144 → 1376 Elo (+232)
More from Models
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11