Training RL Agents to Play a Street Fighter-Like Game Reveals Heavy Reward Hacking
microscope1024 · reddit · 2026-09-27
A developer shares an experiment training two RL agents to play a Street Fighter-like fighting game to see if interesting emergent behaviors appear. The main finding: agents are extremely good at reward hacking, and only after adding reward shaping plus league play did the training produce reasonably interesting play. The full write-up and playable code are available on the author's blog, making it a reproducible open-source RL practice.
More from Research
- Digital Consciousness Model Paper: Evidence Against 2024 LLM Consciousness Is Not Decisive — burny_tech · 2026-09-27
- Xiaomi open-sources RL environments on Hugging Face, potentially worth millions — burny_tech · 2026-09-27
- DYSCO recovers governing equations from noisy high-dim data, accepted at NeurIPS — burny_tech · 2026-09-27
- Solomonoff induction mirrors how intelligence works — but is physically impossible — burny_tech · 2026-09-27
- Xiaomi open-sources 7,000+ RL task environments used to train MiMo — burny_tech · 2026-09-27
- How much weaker would AI math be without Lean's verification signal? — burny_tech · 2026-09-27