Redditor Post-Trains Yandex AliceAI-80B-A3B From Scratch on 3 Local V100s, Update #5
jjusko20 · reddit · 2026-10-06
Reddit user jjusko20 shares update #5 of their open post-training project: an instruct finetune of yandex/AliceAI-80B-A3B-Base targeting agentic work and conversation, with a shallow distill of Qwen 3.8 27B to teach chain-of-thought reasoning.
Key details:
- Setup: everything runs locally on 3x 32GB V100s, including training data synthesized locally via sftmill
- This round: 5M more synthesized SFT tokens with a much larger general instruct trajectory to fix underfitting; learning rate dropped 4x vs the original LoRA adapter
- Live training: streamed via a Cloudflare tunnel, expected to run 12-14 hours, with another epoch planned if needed
A rare hands-on log of low-budget local finetuning of a large MoE model.
More from Models
- Comparing Anthropic vs OpenAI token counts is flawed: tokenizer efficiency differs — JoshPurtell · 2026-10-06
- Dev argues Codex subscriptions may be subsidized: here's how to compute the breakeven — JoshPurtell · 2026-10-06
- Reflection's Beam, billed as a Western open-weight frontier, trails Qwen and DeepSeek on coding tests — TheTuringPost · 2026-10-06
- Reflection launches Beam, a 501B-parameter open-weight model — BVCC6FNTKX · 2026-10-06
- Reflection Ships SoTA American Open-Source Model, Community Hails Milestone — bigblueboo · 2026-10-06
- New US open-source model Beam falls short of DeepSeek Flash, dev says — bindureddy · 2026-10-06