Dev argues DeepSeek V4.1 could have used a non-causal encoder for stronger reasoning

mgostIH · x · 2026-09-11

Developer mgostIH offers an architectural hindsight take: DeepSeek V4.1's architecture could have paired a non-causal encoder with a decoder instead of making both causal — a design he believes would likely yield a much stronger reasoner. The full objective could have trained only on decoded tokens while still remaining causal. This is personal speculation, not official information.

Original post →

More from Models

Models channel →