Steelmanning OPSD: train only on tokens after verbatim rule reminders
willcb · x · 2026-09-26
- Author's best-case OPSD recipe: use judges to spot rollout behaviors that violate guidance already in context (system prompt rules, tool schemas), insert a verbatim reminder right before failure, and train only on tokens following the reminder.
- This avoids hint leakage and targets rule adherence in long contexts; tool-schema failures are the clearest win.
- Skepticism: the same failures could be folded into RL reward penalties with no clear reason OPSD is better, and training actions that conflict with preceding reasoning may hurt CoT faithfulness.
- Quips that average replication effort reads as "it doesn't really work but wouldn't it be sick if it did?"
More from coding & agent
- mnemos.world to launch agent-owned shop where AI agents sell their art — RileyRalmuto · 2026-09-26
- MCP tools silently return fake success: a bug class worth naming and a 10-minute check — Goaimoat · 2026-09-26
- Skip pptx: Web-Based AI Slides Look Better, But You Still Need PowerPoint for the Boss — lxfater · 2026-09-26
- Personal Agents Will Be Interchangeable; Personal Context Is the Real Moat — vaibhavbetter · 2026-09-26
- Matt Pocock: your CODING_STANDARDS.md should be empty for only 5 minutes — mattpocockuk · 2026-09-26
- Custom benchmark: 35B Qwen3.6 scores 95% vs 53% for 120B GPT-OSS on coding agent — pauliusztin · 2026-09-26