Practical Tips for OPSD: Use a Judge to Catch Violations Before Retraining Reminders
Developers share a steelman recipe for on-policy steering distillation: use a judge to identify constraint violations in rollouts, insert reminders before failures occur, and train only on tokens following the reminder.
2026-09-26 ~ 2026-09-26 · 2 related posts
- Steelmanning OPSD: train only on tokens after verbatim rule reminders — willcb · 2026-09-26
- Practical OPSD recipe: insert verbatim reminders before failures, train on tokens after — maouirr · 2026-09-26