Multi-effort RL reward design: R = S − λ_e·C to push up the whole Pareto frontier
carlesgelada · x · 2026-09-12
Carles Gelada shares his team's multi-effort RL setup aiming to raise the entire Pareto frontier: from first principles they derive the reward R = S − λe·C, with S ∈ {0,1} for solve/fail, C the rollout cost, and λe an effort-specific cost penalty. Later thread posts explain how to pick λe.
Related event: RL Reward Design: Choosing λ_e to Lift the Whole Pareto Frontier(3 posts)→
More from Research
- Synthetic Morphology Suggests Non-Physicalist Models of Mind Can Be Empirically Tested — ZeroStateReflex · 2026-09-12
- Apple's Internalized Visual Thinking Drops the Paint-the-Future Pipeline for ~5x Faster Video Reasoning — jiqizhixin · 2026-09-12
- Skild AI founder explains why robotics data needs four sources, each flawed — deepakpathak · 2026-09-12
- Sony CSL's Frank Nielsen releases guaranteed arbitrary-precision approximations of Fisher-Rao geodesic distance — FrnkNlsn · 2026-09-12
- LittleLearner: a 5B model trained from scratch on a K-5-only corpus tests education data limits — repligate · 2026-09-12
- Conjectures launches Bittensor bounties paying TAO for cracking math problems open 30-80 years, judged by machine — markjeffrey · 2026-09-12