Dev post-trains Yandex's 80B model from scratch: first SFT round underfits at 5M tokens
jjusko20 · reddit · 2026-10-05
Redditor jjusko20 gave update #4 on post-training Yandex's open-source AliceAI-80B-A3B-Base into an instruct model for agentic work and conversation, using a shallow distill of Qwen 3.8 27B to teach chain-of-thought.
Results and issues:
- Round 1 SFT completed: the model picks up chain-of-thought and can converse
- But it's badly underfit — the initial 5M tokens were too narrow; ambiguous prompts far from training data produce garbled output
- Checkpoint #1 deemed useless and unreleased
Next steps and tooling:
- Scaled the local synthetic data generator from 80 to 240 tps by pulling from multiple base URLs
- Building a new 5M-token broader general-instruct dataset (previous one was coding-oriented), continuing training at a reduced learning rate
- Open-sourced his off-policy distillation engine sftmill (GitHub), which turns behavioral goals into raw training data
More from Research
- Why FDT is hard: Newcomb variants where the agent simulates a weaker predictor — jessi_cata · 2026-10-05
- Xaira unveils AI drug discovery platform: 10x medicines goal, early results on hard GPCR target — BoWang87 · 2026-10-05
- The Agent Simulates Predictor problem: should one-boxing survive a weaker Omega? — jessi_cata · 2026-10-05
- Bergson: open-source library unifies data attribution methods to study LLM generalization and misalignment — zetalyrae · 2026-10-05
- Aeon classic revisited: your brain does not process information and is not a computer — AnnaCiaunica · 2026-10-05
- The top 50 AI researchers ranked by citations, Attention authors all on the list — ksprdk · 2026-10-05