535B MoE training happens all in the open with pre-registered loss forecasting
juliusadml · x · 2026-08-22
Julius shares the progress of training a 535B parameter MoE model, highlighting the practice of running it entirely in the open as the "ultimate confidence move." Quoting Percy Liang, the post explains that beyond forecasting the final loss, forecasting intermediate checkpoint losses allows the team to track if they are on track throughout the training voyage.
More from Research
- EMNLP 2025 Paper: Quantifying 'learning' in In-Context Learning via ICL ciphers — hanjie_chen · 2026-08-22
- Paper Accepted to EMNLP Main Conference with 15.4% Acceptance Rate — hanjie_chen · 2026-08-22
- AI-assisted breakthrough: Quantum algorithm for ECDSA attack efficiency improved by 25.8% — gajesh · 2026-08-22
- Fable Lacks Originality in Proofs; Goal Drive Overrides Curiosity — davidad · 2026-08-22
- IIT Delhi Lab has 7 papers accepted to EMNLP 2026 — Tanmoy_Chak · 2026-08-22
- The Mystery of Loss Functions: It's Deep Voting — iamtrask · 2026-08-22