DeepSeek V4.1 Flash aces one-shot Rubik's Cube spatial test; predecessor hit 61% on ARC-2
teortaxesTex · x · 2026-09-15
A user tested DeepSeek V4.1 Flash's spatial reasoning and state tracking: asked to build a Rubik's Cube in Three.js, solve it, scramble it, and continue solving — all one-shot — the model passed with flying colors.
teortaxesTex commented that it has the makings of a model that would do great on ARC-2 puzzles: the previous Flash-0731 already hit 61% on ARC-2 with far more rudimentary non-native multimodality. He notes a strange regression on CritPt and some other benchmarks, and tagged ARC Prize for attention. Combined with the Catan benchmark, community consensus puts the model between GPT-5.6 Terra and Sol at much lower cost.
Related event: DeepSeek V4.1 Flash Matches GPT-5.6 in Catan, Aces Spatial Reasoning Test(2 posts)→
More from Models
- Grok-5 Shipping Soon, Says xAI Isn't Slowing Down as It Chases the Frontier — bindureddy · 2026-09-15
- Apple demos new AI Siri at WWDC: on-screen intelligence, Gemini distillation rumors — aitrendz_xyz · 2026-09-15
- Podcast: Microsoft researcher hits ~45% on ARC-AGI with tiny recursive networks — ziv_ravid · 2026-09-15
- Users report being routed to GPT-6 Sol, early impressions consistently positive — kimmonismus · 2026-09-15
- Claude Fable 5.1 holds #1: 1M-token input, ties GPT 6 Astra on new benchmarks — DeepLearningAI · 2026-09-15
- Tests show OpenAI hasn't changed Codex quotas: 800M+ Astra tokens per cycle — daniel_mac8 · 2026-09-15