Delay-corrected Bellman operator + causal attribution for constrained RL
No_Cauliflower7923 · reddit · 2026-08-24
The author proposes CCPL (Causal Consequence-Penalized Learning), tackling the problem in constrained RL where, under stochastic consequence delays, the action that merely preceded a violation gets penalized instead of the causal one.
- Delay-corrected Bellman operator: learns an adaptive effective discount from the consequence-delay distribution, with a contraction proof that holds under unknown stochastic delay.
- Interventional Consequence Net (ICN): pretrained on structural-causal-model labels, estimates marginal causal contribution per action for attribution rather than penalizing by temporal proximity.
- Limitations: the ICN requires access to the environment's structural causal model to generate pretraining labels, so it can't be learned end-to-end from observational or interventional data alone, limiting applicability outside settings where the SCM is known.
The author is looking for collaborators in constrained/safe RL or causal inference.
More from Research
- Task-CoEvolve Cuts Evaluation Cost by 80% via Adaptive Sampling — burny_tech · 2026-08-24
- ToMoE paper converts dense LLMs to MoE without fine-tuning — pmttyji · 2026-08-24
- dots3-note Preview Demonstrates Long-Horizon Agency Capabilities — rohanpaul_ai · 2026-08-24
- dots3-note Preview: 16B Active Parameters Model for Long-Horizon Agency — rohanpaul_ai · 2026-08-24
- Air pollution linked to millions of cancer cases globally: PM2.5 8.8M, NO2 6.7M, ozone 2.6M — EricTopol · 2026-08-24
- DeepSeek V4 Flash solves ARC-AGI task with raw reasoning trace revealed — mhmazur · 2026-08-24