Can LLMs Write Fast Multi-GPU Kernels? Together AI Reveals Performance Gaps

AI Engineer · youtube · 2026-08-28

With NVIDIA's B200 offering 7.2x more compute than the A100 but only 3x better interconnects, the bottleneck in AI workloads has shifted from the GPU to the network links. Together AI introduced ParallelKittens to overlap compute and communication. Benchmarks on ParallelKernelBench show top frontier models solving 28 problems zero-shot, with 31% outperforming baselines. An agent harness raises correctness to 35%, though models still struggle with collective ordering and data partitioning.

Original post →

More from Infra

Infra channel →