AI as compiler: model writes PTX directly, 1.37x speedup on FlashAttention over Triton
Azaliamirh · x · 2026-10-01
New research led by cosfrancois with EPFL's Charly Castes and Thomas Bourgeat proposes 'AI as a compiler': a model translates high-level source (e.g., Triton) directly to low-level assembly (e.g., PTX), with a verifier checking functional correctness via symbolic/numeric analysis plus race conditions and deadlocks — no manually constructed IR/DSL layers. On B200 it beats Triton baselines: FlashAttention 1.37x, Mamba2 state forward 1.10x. This could dissolve the long-standing new-chip software bring-up bottleneck.
More from Infra
- PyTorch Foundation's Mark Collier: open source is the coordination layer for frontier AI — PyTorch · 2026-10-01
- OpenAI and Synopsys sign multi-year deal to build GPT-Synopsys for chip design — BenBajarin · 2026-10-01
- Cognition Becomes First CoreWeave Vera Rubin NVL72 Customer, Sees 4.8X SWE-2 Throughput Boost — altryne · 2026-10-01
- AMD users report 2x faster ComfyUI and no more system lockups after upgrading to ROCm 10 — God_Hand_9764 · 2026-10-01
- Nvidia authorizes $150 billion buyback, a vote on future AI demand — YvesMulkers · 2026-10-01
- GLM 5.3 Flash kernels rewritten on RunInfra: 670 tok/s, 99.7% cache hit, AMD support — ycombinator · 2026-10-01