Feeding Rubin's tensor cores took more than wider MMAs: reuse and deeper pipelines

togethercompute · x · 2026-09-11

Follow-up in the same thread: Rubin's tensor cores can consume operands twice as fast, but simply widening the MMA wasn't enough — the team had to increase data reuse, deepen the pipeline, and more to keep the cores fed, ultimately pushing their 16k NVFP4 GEMM to 22.4 PFLOPS.

Related event: Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS(6 posts)→

Original post →

More from Infra

Infra channel →