Gait Training Tip: Use Rolling Avg Velocity Reward
yacineMTB · x · 2026-09-01
A developer found that tracking immediate velocity rewards caused gait oscillation in RL training. By switching to a reward function based on rolling average velocity over a longer time window, the new policy significantly outperformed the previous trainer.
More from Research
- Dan Luu on why software slowness is a choice, analyzing latency costs and optimization — JeremyCMorgan · 2026-09-01
- Scholar calls out LLM gibberish: reviewing papers and replies is now a waste of time — thegautamkamath · 2026-09-01
- Paper analyzes reasoning models like o1 and DeepSeek R1, probing CoT data contamination — rao2z · 2026-09-01
- Qdrant's Sept 17 stream: token-native storage claims 10-100x faster reads — qdrant_engine · 2026-09-01
- Nature Paper: Interpreting LLM Behavior via Role-Play Framing — mpshanahan · 2026-09-01
- AI Agent Optimization Pressure May Exceed Goodhart's Law — dhadfieldmenell · 2026-09-01