Dataset Policy Gradients: researchers embed a QR code in GPT-2's weights using only training data
ChengleiSi · x · 2026-10-07
A new COLM paper introduces Dataset Policy Gradients, a method to precisely optimize synthetic training data against any differentiable training or post-training metric.
- The approach treats dataset selection itself as an optimizable policy, using policy gradients to steer training data toward a target metric
- Demo: by manipulating training data alone, the authors embedded a scannable QR code into GPT-2's weights, showing precise control over internal representations
- The authors present the work as a poster at COLM
More from Research
- Rubric-conditioned self-distillation for reward supervision accepted at COLM 2026 — armancohan · 2026-10-07
- Jordan and coauthors tackle when to stop generator-verifier loops while controlling false discovery — _onionesque · 2026-10-07
- GroundedSLAM debuts, decisively beating all methods on Meta's egocentric SLAM benchmark — Scobleizer · 2026-10-07
- HCI researcher begs authors to stop claiming 'reflexive' thematic analysis without reflexivity — IanArawjo · 2026-10-07
- Reza Zadeh claims faster matrix multiplication algorithm, suspects labs near exponent 2 — Reza_Zadeh · 2026-10-07
- Redditor proposes graph-based deterministic modeling to make LLM finance agents trustworthy — jonnylegs · 2026-10-07