New Connections Eval Benchmark Ranks Grok Top with High Efficiency
aronchick · x · 2026-08-18
Matson released results for the Connections Eval, a new benchmark based on NYT Connections puzzles. The leaderboard covers 42 models using a one-shot format over 20 games.
Top Results:
- x-ai/grok-4.6: 91 pts ($0.015/game)
- meta/muse-spark-1.1: 90 pts ($0.021/game)
- google/gemini-3.1-pro-preview: 86 pts ($0.021/game)
- openai/gpt-5.6-terra: 86 pts ($0.024/game)
The benchmark focuses on reasoning capabilities in word association puzzles.
More from Models
- DeepSeek harness praised as visionary despite rough edges — aiamblichus · 2026-08-18
- DeepSeek Flash beats Pro on benchmarks with planner-agent workflow — AccBalanced · 2026-08-18
- Reasoning Models Face Persistent Complaints: Opus, Muse, and Gemma — MerePotato · 2026-08-18
- Gemini 3.7 Flash Launches; Box and Databricks Adopt for Real Workflows — DynamicWebPaige · 2026-08-18
- User Reports Codex Burning Through Weekly Quota: 15% in Half a Day — GabGarrett · 2026-08-18
- Anthropic Completes Mythos 2 Training But Declines Release; Mythos 3 Loop Active — kimmonismus · 2026-08-18