Together Compute ships early Rubin GPU kernels, unlocking wider MMA steps and new memory features

togethercompute · x · 2026-09-11

Together Compute published a blog detailing how it optimizes GEMM kernels for NVIDIA's Rubin GPUs: starting from its original B200 GEMMs and progressively adding Rubin-specific features such as wider MMA steps, more tensor and shared memory, B-side collectors, and early A release. The team stresses the kernels are still early, with LUT GEMMs, hardware-native megakernels, and new engine-level optimizations coming next, and credits NVIDIA for early access and support.

Related event: Together AI ports ThunderKittens to Vera Rubin, hits 22 PFLOPS(6 posts)→

Original post →

More from Infra

Infra channel →