Muse Spark 1.1 Leads Medical Benchmarks
alexandr_wang · x · 2026-07-14
Muse Spark 1.1 is hailed as the SOTA on HealthBench Professional.
The cited content compares it to GPT-5.6 Sol: Muse Spark 1.1 scores higher overall. While their length-adjusted scores are statistically similar, Muse Spark 1.1 boasts lower inference costs, with output pricing being roughly 7 times cheaper ($1.25/$4.25 vs $5/$30 per M tokens in/out).
More from Models
- Grok 4.5 is now free inside Cursor, the popular AI coding IDE — mark_k · 2026-07-21
- GPT often converges on the same near-miss ideas in math problems — yacineMTB · 2026-07-21
- Eno Reyes says model distillation is basically unstoppable — LangChain · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- OpenAI hackathon project stalls as Codex struggles on voice, while Claude spots the issue — ColleenMBrady · 2026-07-21
- Kimi K3 lands exactly on China’s 2-year AI capability trend line — peterwildeford · 2026-07-21