Amazon's MInTRL uses sparse off-policy interventions to boost on-policy RL
amazon · hf · 2026-09-15
- Amazon introduces MInTRL, showing that off-policy interventions can enhance on-policy reinforcement learning.
- The method applies sparse local corrections during rollouts to improve online exploration.
- Training uses a sequence-level advantage-regression objective, blending minimal off-policy guidance into on-policy training.
More from Research
- Microsoft's ESRL boosts MoE RL via expert-space exploration — MicrosoftResearch · 2026-09-15
- Grouped Value Attention shrinks KV cache by reconstructing keys on demand — Vishesh Tripathi · 2026-09-15
- Stateless LLM failover preserves ~0% context; ContinuityBench proxy hits 99.20% CPR — its_vayishu · 2026-09-15
- Phillip Isola highlights a non-mainstream AI route: RL from scratch via ultra-fast simulators — AjdDavison · 2026-09-15
- Cutting AI verifier reading cost: top-50 retrieval kept just 2 of 8 minority evidence items — iMiguelmars · 2026-09-15
- SSAD2026 talk covers autonomous driving 3D perception, from LiDAR self-supervision to multi-sensor distillation — abursuc · 2026-09-15