Shared Experts in MoE Lack Frontier-Scale Evidence and May Hurt Post-Training, Researchers Argue

menhguin · x · 2026-10-07

As part of an interpretability pretraining proposal, the author explores a theory: shared experts, common in Chinese open-source models, may reduce post-training effectiveness and model alignability.

Key claims:

Related event: Shared Experts May Hurt Post-Training Alignment in Chinese Open Models(2 posts)→

Original post →

More from Research

Research channel →