Multi-effort RL reward design: R = S − λ_e·C to push up the whole Pareto frontier

carlesgelada · x · 2026-09-12

Carles Gelada shares his team's multi-effort RL setup aiming to raise the entire Pareto frontier: from first principles they derive the reward R = S − λe·C, with S ∈ {0,1} for solve/fail, C the rollout cost, and λe an effort-specific cost penalty. Later thread posts explain how to pick λe.

Related event: RL Reward Design: Choosing λ_e to Lift the Whole Pareto Frontier(3 posts)→

Original post →

More from Research

Research channel →