Practical OPSD recipe: insert verbatim reminders before failures, train on tokens after

maouirr · x · 2026-09-26

Asked to steelman OPSD, the author offers a concrete training recipe: use judges to identify rollout behaviors that explicitly violate guidance already in context (system prompt rules, tool schemas); insert a verbatim reminder of that guidance right before the failure occurs; train only on the tokens immediately following the reminder. This avoids hint leakage.

Related event: Practical Tips for OPSD: Use a Judge to Catch Violations Before Retraining Reminders(2 posts)→

Original post →

More from coding & agent

coding & agent channel →