New model's architecture is 'vanilla': SWA plus MoE with no shared experts, unlike DeepSeek
nrehiew_ · x · 2026-09-23
nrehiew reviews a new model technical report: the architecture is surprisingly simple and vanilla—sliding window attention (SWA) and standard MoE with no shared experts—marking a stark contrast with DeepSeek's approach. He notes the interesting parts are the data and experiments, skipping most architecture details.
More from Research
- Genome LM Minerva finds reverse transcriptase systems encode diverse structured ncRNAs — BrianHie · 2026-09-23
- q-Neurons: stochastic Jackson-derivative activations consistently beat standard ones — FrnkNlsn · 2026-09-23
- OpenAI said to launch journal with multi-agent AI reviews, threatening ML conferences — kfountou · 2026-09-23
- Yale PhD student open-sources his paper figure scripts, packaged as a Skill for Claude Code and Cursor — burny_tech · 2026-09-23
- AI models now match superforecasters on ForecastBench; rematch set for October — burny_tech · 2026-09-23
- Dev uses Opus 5.5 with Lean to formally verify Claude Agent SDK, yielding 16 bug-fix PRs — bcherny · 2026-09-23