USC proposes TOPL, a token-level off-policy method that generalizes across 11 summarization datasets
UniversityofSouthernCalifornia · hf · 2026-07-21
## What TOPL changes Researchers from USC propose **Token-Level Off-Policy Labeling (TOPL)**, an off-policy post-training method that reframes generation learning as a **token-level correctness prediction** problem. Instead of directly training the model to emit off-policy tokens, TOPL teaches it to distinguish good and bad tokens in responses. ## Results - Strong **out-of-distribution generalization** on document summarization across **11 datasets**. - Transfers effectively to **machine translation** as well. - Ablations show the **token-level signal** is crucial; sequence-level analogues do not work as well. - The learned **LoRA adapters** appear interpretable, functioning like linear classification heads and steering vectors.
More from Research
- AlphaFold-guided protein engineering screens 45,000 oxidases and 500 million variants — pushmeet · 2026-07-21
- OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans — Baochen Fu · 2026-07-21
- LTX-2.3 face-and-voice LoRA training can work on 12GB VRAM with heavy tradeoffs — __alpha_____ · 2026-07-21
- Follow-up paper argues digital twins could make clinical trials more adaptive — techhalla · 2026-07-21
- Nature npj Digital Medicine paper maps causal inference and digital twins for trials — techhalla · 2026-07-21
- AI performance is increasingly limited by materials science, not just compute — nordicinst · 2026-07-21