DeepSeek V4.1 'humiliates' rival in creative game design, RL approach hailed as vindicated
teortaxesTex · x · 2026-09-09
Blogger teortaxesTex says he's amazed by DeepSeek V4.1 regardless of benchmarks, calling it total vindication of DeepSeek's RL research program with R1-Zero energy for the agentic era: the model plays freely like a child with no fear, malice, or visible alignment constraints, and he argues DeepSeek should ignore competition and scale this approach. He adds that V4.1 (Neowhale) is outperforming another model (Astra) in creative game design.
More from Models
- Watch Astra beat Montezuma's Revenge, RL's infamous benchmark game — emollick · 2026-09-09
- Model Can't Draw ASCII Whales, So It Writes a Node.js Script to Improve — teortaxesTex · 2026-09-09
- Blender head-to-head: same prompt, 12 seconds, and one frontier model is in a different class — ZeroStateReflex · 2026-09-09
- Netlify adds GPT-6 Astra, Gemini 3.8 Flash, Claude Fable 5.1 and smarter Agent Runner scoping — thisiskp_ · 2026-09-09
- GLM-5.3-Flash tops agentic tool-call leaderboard at 78%, priced at just $0.50/M output tokens — shensi · 2026-09-09
- Artificial Analysis Updated Benchmarks Twice in 4 Days for Astra — py-net · 2026-09-09