RLTL;DR paper: self-generated feedback breaks Pass@128=0 barrier to 14-31% Pass@1

burny_tech · x · 2026-10-01

Michael Kirchhof and colleagues introduce RLTL;DR, a self-improvement method for reinforcement learning with verifiable rewards. When tasks are so hard that agents fail all 128 attempts (Pass@128=0), standard GRPO has no signal to optimize. RLTL;DR instead has the policy write a one-line TL;DR insight after each failed attempt, conditioned on all prior insights, and backpropagates through the in-context insights to internalize a task-to-insight mapping. On tool-calling and coding datasets where Qwen 3.5 9B Thinking stays flat at 0-1% Pass@1 under GRPO, RLTL;DR reaches 14-31% during training and, crucially, 12-13% at eval with no insights in context. The authors also distill the approach into SFTL;DR, training only on (task, insight) tuples, to study the internalization mechanism.

Original post →

More from coding & agent

coding & agent channel →