Dev argues DeepSeek V4.1 could have used a non-causal encoder for stronger reasoning
mgostIH · x · 2026-09-11
Developer mgostIH offers an architectural hindsight take: DeepSeek V4.1's architecture could have paired a non-causal encoder with a decoder instead of making both causal — a design he believes would likely yield a much stronger reasoner. The full objective could have trained only on decoded tokens while still remaining causal. This is personal speculation, not official information.
More from Models
- V4.1 Ranks #5 on LiveBench, Tops Agentic Coding but Called Language-Skewed — teortaxesTex · 2026-09-11
- OpenAI ships GPT-Live prompting guide: copying your old prompts won't hit SOTA — craigsdennis · 2026-09-11
- GPT-6 Astra burns through user's weekly usage cap, forcing a wait until Monday — max_paperclips · 2026-09-11
- Wiz Launches Cyber Model Arena: Gemini 3.8 Flash Cyber Tops at 74.9% — rseroter · 2026-09-11
- Anthropic blocks minors from using Claude, HN debates age policy — petrusenko_max · 2026-09-11
- V4.1 hits #5 on LiveBench, tops Agentic Coding by 20 points over Astra — teortaxesTex · 2026-09-11