Muse Spark 1.1 Leads in Benchmark Performance
alexandr_wang · x · 2026-07-13
Alexandr Wang shared new evaluation results from theoretical computer science/finite model theory: Muse Spark 1.1 outperforms Opus, Grok 4.5, and Gemini on this benchmark.
The cited benchmark specifies that models are given several small graphs and must output a first-order logic formula describing the properties of designated nodes across multiple graphs simultaneously. The evaluation consists of 64 questions, heavily focusing on inductive reasoning and formal expression capabilities.
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11