Practical OPSD recipe: insert verbatim reminders before failures, train on tokens after
maouirr · x · 2026-09-26
Asked to steelman OPSD, the author offers a concrete training recipe: use judges to identify rollout behaviors that explicitly violate guidance already in context (system prompt rules, tool schemas); insert a verbatim reminder of that guidance right before the failure occurs; train only on the tokens immediately following the reminder. This avoids hint leakage.
More from coding & agent
- KoboldCpp ships built-in Agent harness; author warns of phishing site koboldcpp.com — HadesThrowaway · 2026-09-26
- Nace launches Drex, a sub-6B decision model outputting option probabilities instead of prose — rohanpaul_ai · 2026-09-26
- The team was already running LLM cross-checks on reports — the real gap was that it all lived in one person's account — beglen · 2026-09-26
- Muse's connector system wins praise: one-tap context access to Steam, Bilibili, DeepSeek with keys never exposed to the model — op7418 · 2026-09-26
- Testing Opus 5.5 to generate a promo video from repo code — vista8 · 2026-09-26
- MiniMax H3 long-video editor stitches Latents for seamless 45s+ one-take videos — InariKirin · 2026-09-26