Solo dev trains 3.87B MoE from scratch on just 86.5B tokens, matches Qwen2.5-1.5B on code

Prestigious-Taste-63 · reddit · 2026-10-04

Apex-2: A MoE small model pretrained from scratch by a solo dev

The author trained a decoder-only MoE model Apex-2 entirely from scratch (no external base weights) and open-sourced it on Hugging Face:

Benchmark results (SFT)

Notable: with 0.087T pretrain tokens, the base model's HumanEval+ matched Qwen2.5-1.5B (trained on 18T); knowledge and math still lag due to the data gap.

What didn't work

Original post →

More from Research

Research channel →