USC proposes TOPL, a token-level off-policy method that generalizes across 11 summarization datasets

UniversityofSouthernCalifornia · hf · 2026-07-21

## What TOPL changes Researchers from USC propose **Token-Level Off-Policy Labeling (TOPL)**, an off-policy post-training method that reframes generation learning as a **token-level correctness prediction** problem. Instead of directly training the model to emit off-policy tokens, TOPL teaches it to distinguish good and bad tokens in responses. ## Results - Strong **out-of-distribution generalization** on document summarization across **11 datasets**. - Transfers effectively to **machine translation** as well. - Ablations show the **token-level signal** is crucial; sequence-level analogues do not work as well. - The learned **LoRA adapters** appear interpretable, functioning like linear classification heads and steering vectors.

Original post →

More from Research

Research channel →