NeurIPS paper: value-based RL agents can fail to converge to any optimal policy in Newcomblike environments
jessi_cata · x · 2026-09-16
Cited in a debate on CDT/EDT and RL, Bell, Linsefors, Oesterheld & Skalse's NeurIPS 2021 paper studies value-based RL in Newcomblike environments. Key results: RL agents cannot converge to non-ratifiable policies; some Newcomblike environments admit no optimal policy the agent can converge to; ratifiable policies always exist but convergence is not guaranteed; and several alternative limit behaviors are characterized. The poster notes the paper tests GRPO while theory points to CDT self-ratification, suggesting wrapping self-ratification into policy optimization via policy-dependent KKT conditions.
More from Research
- Composing Continual Learning Mechanisms Boosts Long-Horizon Memorization in LMs — JohnsHopkins · 2026-09-16
- Gavel Elicits Native Skill Routing from a Frozen LLM's Hidden States — Tsinghua · 2026-09-16
- ScienceBuddy: Recursive-in-Recursive Self-Improvement for Scientific Agents — Shuhan Xue · 2026-09-16
- RCT: LLMs fail to significantly boost novices' wet-lab molecular biology success — shae_mcl · 2026-09-16
- Six years on, scvi-tools still widely used — outlasting foundation models and coding agents — anshulkundaje · 2026-09-16
- Lightning Weave composes reasoning capabilities via on-policy distillation — Yecheng Wu · 2026-09-16