Multi-effort RL: set cost penalty to the Pareto curve slope to lift the whole frontier
carlesgelada · x · 2026-09-12
Carles Gelada details a multi-effort RL reward design aimed at pushing up the entire Pareto frontier: reward R = S − λe·C, where S is solve/fail, C the rollout cost, and λe an effort-specific penalty. The key result: the correct λe equals the slope of the base model's Pareto curve at each effort level, approximated once and held constant during RL. This works because λe sets the slope of iso-reward lines; only when those lines are tangent to the Pareto frontier does improving reward guarantee frontier improvement.
Related event: RL Reward Design: Choosing λ_e to Lift the Whole Pareto Frontier(3 posts)→
More from Research
- Tao and Fields Medalists' two objections to AI in math, and why they're weak — RexDouglass · 2026-09-12
- Conjectures launches Bittensor bounties paying TAO for cracking math problems open 30-80 years, judged by machine — markjeffrey · 2026-09-12
- FADA (CoRL 2026) open-sourced: humanoid robots adapt to new conditions from 2 minutes of experience — GuanyaShi · 2026-09-12
- Intel's silicon photonics couplers hit 1-1.5 dB IL, with visible epoxy delamination flaws — jwt0625 · 2026-09-12
- CPO paper criticized for vague DLW-to-PIC coupling description: 'such as TCB' — jwt0625 · 2026-09-12
- Fly connectome trained to play a Flappy Bird–style game — TinfoilTricorn · 2026-09-12