Digimon AI Project: Reward Hacking and Progress in PPO Training

redfoxkiller · reddit · 2026-09-02

The author shares an update on their Digimon AI project (Digi-Brain), which uses PPO to train agents in procedurally generated dungeons. The main challenge is PPO update spikes causing the agent (Yozoramon) to get stuck in local optima (e.g., obsessing over floor 1 cake). The new version improves early learning; while some runs still end on floor 1, stability in reaching floors 3 and 4 has increased. The goal is stable learning without months of training time, moving next to a Neural-MMO.

Original post →

More from Embodied

Embodied channel →