KKT Points in Imperfect-Recall Games Correspond to CDT+GT Optimal Policies

jessi_cata · x · 2026-10-10

jessicata draws out an implication of the IJCAI-23 paper The Computational Complexity of Single-Player Imperfect-Recall Games (Tewolde, Oesterheld, Conitzer, Goldberg): KKT points in imperfect-recall games (e.g., Sleeping Beauty, Absentminded Driver) correspond to CDT+GT optimal policies, and KKT points are stable under ideal policy-gradient optimization. The author notes complexities in connecting these KKT points to optimizing neural-network parameters for memoryless POMDPs, but argues they should align when the network is over-parameterized relative to the policy set. The thread also traces earlier work like Piccione & Rubinstein (1997).

Original post →

More from AGI Musings

AGI Musings channel →