TritonRL: an 8B RL-trained model matches 100B+ frontier models on Triton kernel generation

allenainie · x · 2026-10-09

The paper "TritonRL: Training LLMs to Think and Code Triton Without Cheating" (Jiin Woo, Allen Nie, et al.) introduces a domain-specialized 8B model for Triton GPU kernel programming. Key contributions: a multi-layered verification system providing high-fidelity rewards to guard against reward hacking, and Hierarchical Reward Decomposition (HRD), which decouples reinforcement for high-level reasoning versus low-level implementation to solve credit assignment in long-sequence generation. On KernelBench, TritonRL achieves state-of-the-art correctness and runtime speedup, beating concurrent Triton-specific models and matching 100B+ parameter frontier models—highlighting hardware-aware RL as a cost-effective adaptation paradigm. The authors present at COLM, poster #86.

Original post →

More from coding & agent

coding & agent channel →