14-year-old trains RL agents to race in PolyTrack, reveals extreme closed-loop sensitivity
michael201110 · reddit · 2026-10-06
A 14-year-old developer shared PolyBot, a reinforcement-learning project training AI to drive the racing game PolyTrack, experimenting with TQC, PPO, and GRTQC.
Key findings
- At high speed the system becomes extremely sensitive: a persistent steering offset of just +0.0001 on a trained 24.263s TQC policy caused failures across 56% of the track; ±0.0001 perturbations to longitudinal input could fail before 24% progress — a "racing butterfly effect."
- This led to work on behavioral cloning/teacher-student transfer, reward shaping, curriculum learning, deterministic evaluation, and quantile critics.
Engineering
- PolyBot drives the game directly via a PolyModLoader bridge with fixed physics steps for deterministic training/eval.
- The from-scratch GRTQC agent uses two 128-unit gated networks, twin quantile critics, 121 inputs, and its own curriculum.
Video and open-source code are available; the author seeks feedback on closed-loop sensitivity and plans to visualize hundreds of past training runs as in-game ghosts.
More from Fun
- Garry Tan ports Doom to Paul Graham's Bel LISP in 20 minutes using Opus 5.5 — garrytan · 2026-10-07
- Jon Stewart Roasts Meta's Muse AI Agent; Alexandr Wang Fires Back on X — alexandr_wang · 2026-10-07
- Neil deGrasse Tyson says AI 'knows everything, understands nothing' — and gets mocked — StewartalsopIII · 2026-10-07
- Rumor of an OpenAI math-proof treasure trove? This thread says don't hold your breath — patience_cave · 2026-10-07
- 'Beautiful higher-dimensional structures': an AI-proofs quip becomes an X meme — yacinelearning · 2026-10-07
- 60 builders to voice-enable a Furby at SF Tech Week hardware hackathon — AssemblyAI · 2026-10-07