AI as compiler: model writes PTX directly, 1.37x speedup on FlashAttention over Triton

Azaliamirh · x · 2026-10-01

New research led by cosfrancois with EPFL's Charly Castes and Thomas Bourgeat proposes 'AI as a compiler': a model translates high-level source (e.g., Triton) directly to low-level assembly (e.g., PTX), with a verifier checking functional correctness via symbolic/numeric analysis plus race conditions and deadlocks — no manually constructed IR/DSL layers. On B200 it beats Triton baselines: FlashAttention 1.37x, Mamba2 state forward 1.10x. This could dissolve the long-standing new-chip software bring-up bottleneck.

Original post →

More from Infra

Infra channel →