Discussion on Multi-Agent Reward Schemes and Convergence

jessi_cata · x · 2026-08-27

The author discusses details on multi-agent training, specifically questioning whether agents run on shared reward signals, individual rewards, or a sum. If the agent reward is purely collective, it resembles a single-player game of imperfect recall. The author guesses convergence to CDT+GT or "modified multiself equilibrium", partly based on the general convergence of RL to CDT.

Original post →

More from Research

Research channel →