Moonshot Criticized for Ignoring ARC-AGI Benchmark, Focusing on Coding and Writing
bookwormengr · x · 2026-08-01
A user criticized Moonshot for neglecting challenging reasoning benchmarks like ARC-AGI, arguing the company only focuses on coding, GDPEval, and creative writing. Citing OpenAI as an example, the post noted that while early GPT-5-Pro struggled with ARC-AGI, later versions improved significantly, urging frontier labs to prioritize such revolutionary evaluations. The original post also complained about Kimi's current pricing lacking competitiveness.
More from Models
- Testing Frontier Math Skills: Proving Theorems Harder Than Finding Counterexamples — doodlestein · 2026-08-01
- Lamenting Claude Haiku 3.5: Developers Urge AI Companies to Stop Deprecating Old Models — repligate · 2026-08-01
- Redditor: 30B parameters is the sweet spot for local video models — RekTek4 · 2026-08-01
- Redditor: 30B parameters is the sweet spot for local video models — RekTek4 · 2026-08-01
- DeepSeek Matches Claude Sonnet in Agentic Loops at 600% Lower Cost — bindureddy · 2026-08-01
- Kimi K3 Hits Record 172 Tokens/sec in Inference Speed — AccBalanced · 2026-08-01