Ai2's tech report shows how 512 GPUs train one model—and flags 'token gerrymandering'

allen_ai · x · 2026-10-01

Alongside Olmo-core 3, Ai2 published a tech report and an interactive explainer walking through how 512 GPUs train a single large model: from single-GPU compute/memory basics through data parallelism, memory sharding, and expert parallelism. The report also details experiments and failures, including 'token gerrymandering'—a case where a load-balancing score improved even as workloads became less balanced, showing the metric can be misleading.

Original post →

More from Infra

Infra channel →