Google Cloud reproduces Olmo 3 7B pre-training on TPUs, matching Ai2 on held-out evals
allen_ai · x · 2026-09-29
Google Cloud engineers reproduced Ai2's fully open Olmo 3 7B from scratch using MaxText on Cloud TPUs, covering both stage-1 pre-training and stage-2 mid-training anneal, and verified the match on held-out evals rather than just loss curves.
- PyTorch→JAX conversion validated via logit-parity: KL ≈ 1.5e-3 vs the HuggingFace reference, with 98.75% top-1 token agreement at full 8192-token context in bfloat16.
- Held-out evals caught a data-loader bug that made MaxText look better than the reference — the gain was memorization.
- Olmo 3 was chosen for exposing data, code, configs, checkpoints, logs and evals, plus an independent PyTorch/GPU reference.
A strong case for full openness enabling reproducibility science and trustworthy AI.
More from Infra
- Celesto Launches GitHub Actions Runners, Claims 12x Cheaper Than GitHub — aniketmaurya · 2026-09-30
- Wasmer's Pi runs AI agents unmodified in the browser and on iPhone via WebAssembly — JosephJacks_ · 2026-09-30
- Qualcomm's Kedar Kondap on X2 Elite: new Surface devices, Linux support and the PC landscape — ryanshrout · 2026-09-30
- Book-Length Deep Dive Explains Virtual Memory From First Principles: Page Tables, TLBs, NUMA — abhi9u · 2026-09-30
- Swift 1.5 matches Qwen3.8 27B quality on M5 Max while writing 34% fewer tokens — DerTomsn · 2026-09-30
- Anthropic inks up to $84.5B compute deal with SpaceX through 2029 — XFreeze · 2026-09-30