Muse Spark 1.1 Leads in Benchmark Performance
alexandr_wang · x · 2026-07-13
Alexandr Wang shared new evaluation results from theoretical computer science/finite model theory: Muse Spark 1.1 outperforms Opus, Grok 4.5, and Gemini on this benchmark.
The cited benchmark specifies that models are given several small graphs and must output a first-order logic formula describing the properties of designated nodes across multiple graphs simultaneously. The evaluation consists of 64 questions, heavily focusing on inductive reasoning and formal expression capabilities.
More from Models
- Kimi K3 rises to No. 4 on the Agent Arena leaderboard — HeyZoyaKhan · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22