Salesforce's RISE: Dense Token-Level Supervision From the Model's Own RL Trajectories

Salesforce · hf · 2026-09-07

Salesforce introduces RISE (Recursive Improvement via Self-Extrapolating Policy Distillation), a post-training method that recursively generates dense token-level supervision from the model's own RL trajectory via self-extrapolation.

The approach avoids external teachers or human annotation, offering a path to cheaper post-training by having the model bootstrap its own supervision signal.

Original post →

More from Research

Research channel →