Post-training Yandex AliceAI-80B from scratch: GGUF release and V100 kernel planned
jjusko20 · reddit · 2026-10-02
Reddit user jjusko20 shares progress on post-training Yandex's open-source AliceAI-80B-A3B base into an instruct model:
- The approach is an SFT on a self-built synthetic distilled dataset; training is nearly done. The loss curve looks wacky (vocabulary size, varying epoch token counts), but average per-epoch loss is steadily decreasing.
- Within days they plan to release GGUF quants plus a llama.cpp patch for local inference, and will publish the checkpoint publicly regardless of quality ("kinda like how deepseek did it"), followed by RL and an extended SFT set.
- Already open-sourced the distillation engine (SFTMill), and has a custom training kernel for V100s plus other patches to share.
- The author is also job-hunting for remote/NYC dev or ML engineer roles.
More from Models
- Polymarket Puts Just 10% Odds on OpenAI Fully Pausing AI Training This Year — Polymarket · 2026-10-02
- Apple Intelligence's overnight cron is randomly waking macOS 27 machines — DanielLockyer · 2026-10-02
- EMNLP Findings paper: explicit reasoning hurts pointwise rerankers, disabling it helps — lintool · 2026-10-02
- User calls Claude 4.7 'neurotic': it challenges you over the slightest risk — repligate · 2026-10-02
- JevBench to add evals for LLM routing, RAG retrieval, and moderation use cases — airesearch12 · 2026-10-02
- Opus 5.5's Writing Tell: 'Dependable' Appears 23x More Often Than in Human Text — TechCrunch AI · 2026-10-02