Why Parallax Routed Experts Don't Degrade
jon_durbin · x · 2026-07-16
This discusses a common question about Parallax: why do routed experts only see about 1/G of the tokens without almost any drop in quality?
The author explains that this isn't "magic", but the result of combining several designs:
- Using latent MoE to make individual experts smaller, then adding more experts.
- Increasing top-k so more tokens hit local experts.
- Adding control variate to the router to make routing credit assignment more stable.
The conclusion is that as long as the model is designed "decentrally" rather than forcibly splitting apart a centralized model, the parameters can still obtain a sufficient token budget.
More from Research
- Project APE launches CRED to test whether LLMs can verify research errors — soumitrashukla9 · 2026-07-22
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22