ARC Prize to Evaluate DeepSeek V4.1 Flash After Predecessor Hit 61.4% on ARC-AGI-2
teortaxesTex · x · 2026-09-30
ARC Prize confirmed it will evaluate DeepSeek V4.1 Flash. The predecessor, V4-Flash-0731, scored 61.4% on ARC-AGI-2.
@teortaxesTex notes that two months later, V4.1 is much larger, more advanced, natively multimodal, with 29% more tokens — and Dots3-preview already hits 77%. He argues anything under 80% would be disappointing. ARC Prize replied 'Sooooon.'
Related event: ARC Prize to Benchmark DeepSeek V4.1 Flash(2 posts)→
More from Models
- User claims GPT-6.1 Sol ULTRA ran 25 minutes on just 1% of weekly quota (unverified) — steipete · 2026-09-30
- GPT 6.1 Sol launches with 50% cheaper caching; Luna can also make motion videos — oran_ge · 2026-09-30
- Users report Grok Bot is now nearly as fast as Muse — yunta_tsai · 2026-09-30
- banteg: model recovers 88 decompiled functions per hour that recompile to exact original machine code — banteg · 2026-09-30
- Will AI subscriptions hit $1,000/month next year? Priced against a $200k salary, some say it's cheap — RachelVT42 · 2026-09-30
- Early user on Grok's Dot Bot: 'You can immediately feel the difference' — jdjohnson · 2026-09-30