Orion 16B hits 100B training tokens using DPP on distributed GPUs
markjeffrey · x · 2026-08-18
- Milestone: Orion 16B model has surpassed 100 billion training tokens, with performance still improving.
- Architecture: It is the largest LLM pretrained using DPP to date.
- Scale: The run has reached $10^{22}$ FLOPS, utilizing heterogeneous commodity GPUs from permissionless, globally distributed providers.
- Plan: The team plans to add new GPU types next week, scale to 256 GPUs, and optimize throughput for the next stage.
More from Infra
- SGLang updates Qwen3.8-27B recipes, hitting 206 tok/s on RTX 5090 — ying11231 · 2026-08-18
- Reranking Paradox: Performance Drops as Document Count Increases — CShorten30 · 2026-08-18
- Running Qwen 3.8 27B on RTX 3090: Configuration Guide — cezarducatti · 2026-08-18
- Netlify launches Git host 'Source', claims 2x speed over GitHub — thisiskp_ · 2026-08-18
- Google Reportedly Bidding $10M for Spirit Airlines' Enterprise Data — soumitrashukla9 · 2026-08-18
- Optimizing AI Infra: 4 Core Strategies to Reduce Data Movement — prateekj · 2026-08-18