Researcher: shared experts in Chinese open-source models may hurt post-training alignability
menhguin · x · 2026-10-07
In an interp discussion, the author shares that one theory in their interpretability pretraining proposal is that shared experts — common in Chinese open-source MoE models — may reduce the effectiveness and model alignability of post-training. They traced the lineage of shared experts and found no conclusive evidence at frontier scale that they benefit pretraining. The peer jokes that the multi-token prediction paper still haunts them: every mediocre small-scale idea now tempts them to run more ablations hoping it magically works.
Related event: Shared Experts May Hurt Post-Training Alignment in Chinese Open Models(2 posts)→
More from Models
- Grok 4.7 lands on Microsoft Foundry, expanding xAI's enterprise reach — XFreeze · 2026-10-07
- WebDev Arena: claude-opus-5.5-max tops leaderboard at 1814 with 846K votes cast — arena · 2026-10-07
- Mistral Large 4 Beats Opus 5.5 and GPT-6 Astra on Cybersecurity Benchmarks — Lower Refusal Rates — burny_tech · 2026-10-07
- Independent model Auro V9 ships, beating V8 head-to-head at a 10:1 clip — TheMoonMidas · 2026-10-07
- _xjdr: only four open-weights models are currently interesting — _xjdr · 2026-10-07
- Researcher _xjdr Praises Inkling Architecture, Calls DSV4 Models an Acquired Taste — _xjdr · 2026-10-07