Solo dev trains 3.87B MoE from scratch on just 86.5B tokens, matches Qwen2.5-1.5B on code
Prestigious-Taste-63 · reddit · 2026-10-04
Apex-2: A MoE small model pretrained from scratch by a solo dev
The author trained a decoder-only MoE model Apex-2 entirely from scratch (no external base weights) and open-sourced it on Hugging Face:
- Architecture: every layer is MoE (no dense layers), 3.87B total / 1.45B active params, 32 layers, dmodel 2048, GQA 16Q/4KV, 16 experts top-4, 4k context, Qwen3 tokenizer (151k); loads via transformers / vLLM with the Qwen3MoeForCausalLM mapping
- Training: only 86.5B pretrain tokens (1→2 GH200, using DiLoCo); SFT on 2.5B tokens (code-heavy + math + instruction)
Benchmark results (SFT)
- HumanEval 43.9 / HumanEval+ 41.5
- MBPP 56.3 / MBPP+ 48.9
- GSM8K 32.4, MATH-500 21.0, IFEval 44.7, MMLU 28.6
Notable: with 0.087T pretrain tokens, the base model's HumanEval+ matched Qwen2.5-1.5B (trained on 18T); knowledge and math still lag due to the data gap.
What didn't work
- DPO backfired: 220k length-normalized pairs made answers much longer and hurt code/math/IFEval, so the checkpoint was dropped in favor of SFT
- Honest limitations: English-centric, weak knowledge with hallucinations, near-zero on LiveCodeBench medium/hard, 4k context only
More from Research
- Functionalists vs Landauer: why MoE sparsity analogies may extend to brains and consciousness — JoshPurtell · 2026-10-06
- Updated preprint: language models mirror human content effects on reasoning tasks — AndrewLampinen · 2026-10-06
- Stanford researchers: data augmentation as a lens on hippocampal generalization — AndrewLampinen · 2026-10-06
- DeepMind's Andrew Lampinen: where AI and brains converge on similar solutions reveals invariant computational principles — AndrewLampinen · 2026-10-06
- Virtual cell debate was a metrics problem: Nature Biotech paper says calibrated metrics vindicate perturbation models — anshulkundaje · 2026-10-06
- Guava open-sources VLM distillation: 2K sim trajectories yield a 4B robot agent at 90% real-world success — furongh · 2026-10-06