Digimon AI Project: Reward Hacking and Progress in PPO Training
redfoxkiller · reddit · 2026-09-02
The author shares an update on their Digimon AI project (Digi-Brain), which uses PPO to train agents in procedurally generated dungeons. The main challenge is PPO update spikes causing the agent (Yozoramon) to get stuck in local optima (e.g., obsessing over floor 1 cake). The new version improves early learning; while some runs still end on floor 1, stability in reaching floors 3 and 4 has increased. The goal is stable learning without months of training time, moving next to a Neural-MMO.
More from Embodied
- Expert: 1 million humanoids in US jobs within a decade, maybe — binarybits · 2026-09-02
- ZimaBlue: Evolving Generalizable World Action Models via Video Pre-training — JoyFutureAcademy · 2026-09-02
- Qwen-Drive-1.0: A Vision-Language Foundation Model for Autonomous Driving — Qwen · 2026-09-02
- Analyst report massively overestimates robot data generation — zephyr_z9 · 2026-09-02
- Markov Robotics demos sub-millimeter pick and place, highlighting zero-shot generalization for physical AGI — Scobleizer · 2026-09-02
- NVIDIA Warp hits 10M downloads; livestream to cover simulation and robotics workflows — milesmacklin · 2026-09-02