Muse Spark 1.1 Leads Medical Benchmark
Daisy4ai · x · 2026-07-14
On HealthBench Pro (OpenAI's benchmark of 525 real-world clinical tasks), Muse Spark 1.1 performs similarly to GPT-5.6 Sol, with an overall score that is even slightly higher.
The original post highlights that its length-adjusted score is basically on par with GPT-5.6 Sol, but at a much lower reasoning cost: input/output is about $1.25/$4.25 per million tokens, compared to $5/$30 for GPT-5.6 Sol, making the output side roughly 7 times cheaper. The author sets "affordable medical superintelligence" as the ultimate goal.
More from Models
- 6TB of Fable data sold with leaked SSH keys, cloud creds tied to Xiaomi, Huawei, NIO — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11