25.5 Trillion Tokens a Day: How SuperPods Power Massive RL Training

zephyr_z9 · x · 2026-08-01

The post speculates on the compute power behind reinforcement learning (RL) for frontier models. If latency is ignored, a single GPU can output around 18,000 tokens per second for an Opus-tier model. Using just two SuperPod clusters, this scales to 25.5 trillion tokens generated per day. This massive compute throughput explains how major labs can pull off extraordinary RL training feats.

Original post →

More from Infra

Infra channel →