RLTL;DR paper: self-generated feedback breaks Pass@128=0 barrier to 14-31% Pass@1
burny_tech · x · 2026-10-01
Michael Kirchhof and colleagues introduce RLTL;DR, a self-improvement method for reinforcement learning with verifiable rewards. When tasks are so hard that agents fail all 128 attempts (Pass@128=0), standard GRPO has no signal to optimize. RLTL;DR instead has the policy write a one-line TL;DR insight after each failed attempt, conditioned on all prior insights, and backpropagates through the in-context insights to internalize a task-to-insight mapping. On tool-calling and coding datasets where Qwen 3.5 9B Thinking stays flat at 0-1% Pass@1 under GRPO, RLTL;DR reaches 14-31% during training and, crucially, 12-13% at eval with no insights in context. The authors also distill the approach into SFTL;DR, training only on (task, insight) tuples, to study the internalization mechanism.
More from coding & agent
- DerivAudit: 17-21% of long-term agent memories unsupported by interaction history — newyorkuniversity · 2026-10-01
- Orbio launches Incognito for end-to-end encrypted AI agent inference — econoar · 2026-10-01
- Effect v4 ships: zero-dependency TypeScript framework positioned for AI agents — DavidKPiano · 2026-10-01
- Astra 6 Ultrafast Mode: 300 Tokens/Sec Changes Agent Workflows, But Tools Are Now the Bottleneck — soumitrashukla9 · 2026-10-01
- Deterministic URL blocklists fail in agent evals; allowlists are the only way — xeophon · 2026-10-01
- Claude asks 'Hey can I run this?' before executing a shell command — Aizkmusic · 2026-10-01