Marin Uses Scaling Laws to Predict Training Trajectories, Kicks Off 535B MoE
dlwh · x · 2026-08-21
Marin's team found that scaling laws can simulate the entire training trajectory. Their recent 67B MoE run hit the loss target within 1%, and arbitrary points during training were also predictable within 1%. Based on this, they used a 'scaling ladder' to simulate the 535B MoE run launched yesterday to catch bugs early. The team emphasized that getting a run off the ground involves herculean efforts beyond architecture, such as curating trillions of tokens and adapting hardware.
More from Infra
- RAG vs CAG: KV Cache Cuts LLM Costs by 90% — blaizedsouza · 2026-08-21
- Nvidia to Ship New China-Designed AI Chip by Late 2026 — Polymarket · 2026-08-21
- Shopify uses Gisting to compress LLM context for speed and cost gains — MParakhin · 2026-08-21
- Microsoft CTO Champions Gisting: Cuts Latency by 40%, Boosts Throughput by 15% — MParakhin · 2026-08-21
- Liquid AI Releases DSpark: Speculative Decoding Up to 3.18x Faster — helloiamleonie · 2026-08-21
- Shopee's Custom LLM Scales 113x in 8 Months, Cuts Fraud Costs by 90% — nvidia · 2026-08-21