Delay-corrected Bellman operator + causal attribution for constrained RL

No_Cauliflower7923 · reddit · 2026-08-24

The author proposes CCPL (Causal Consequence-Penalized Learning), tackling the problem in constrained RL where, under stochastic consequence delays, the action that merely preceded a violation gets penalized instead of the causal one.

The author is looking for collaborators in constrained/safe RL or causal inference.

Original post →

More from Research

Research channel →