Post-training Yandex AliceAI-80B-A3B from scratch: a NaN bug in custom V100 kernels killed one run
jjusko20 · reddit · 2026-10-03
An update on the author's from-scratch post-training of Yandex's AliceAI-80B-A3B (instruct):
- The previous run was cooked: after evaluating the QLoRA, all layers but two had zero gradients — a NaN issue from the author's custom V100 kernels. Two layers alone emulated a plausible loss curve until evaluation exposed it.
- Fixed and restarted, with training livestreamed again via a Cloudflare tunnel; the new loss curve looks much healthier.
- The post includes first-epoch loss curve comparisons from both runs.
More from Models
- Open-source project clef hits #3 on Hugging Face trending — michellechen · 2026-10-03
- Arena Weekly: Gemini 4 Argon Tops Text Arena, Sonnet 5.5 Within 2 Points of GPT-6 at 80% Less Cost — arena · 2026-10-03
- OpenTumorBoard: 611 real tumor board cases show frontier LLMs still lag specialists — Anqi Li · 2026-10-03
- Keyword benchmarks credit fake tool use in small models; a 3.3 GPU-hour SFT repairs it — Juan S. Santillana · 2026-10-03
- NVIDIA fine-tunes Nemotron ASR for Saudi dialects, cutting WER from 55% to 30% — NVIDIAAI · 2026-10-03
- Early tip: set Opus 5.5 to medium — prasenx · 2026-10-03