Researcher: shared experts in Chinese open-source models may hurt post-training alignability

menhguin · x · 2026-10-07

In an interp discussion, the author shares that one theory in their interpretability pretraining proposal is that shared experts — common in Chinese open-source MoE models — may reduce the effectiveness and model alignability of post-training. They traced the lineage of shared experts and found no conclusive evidence at frontier scale that they benefit pretraining. The peer jokes that the multi-token prediction paper still haunts them: every mediocre small-scale idea now tempts them to run more ablations hoping it magically works.

Related event: Shared Experts May Hurt Post-Training Alignment in Chinese Open Models(2 posts)→

Original post →

More from Models

Models channel →