Thermal-aware RL walks Unitree A1 27 min under 3 kg; baseline overheats in 7

Learning Thermal-Aware Locomotion Policies for an Electrically-Actuated Quadruped Robot

Letian Qian, Yuhang Wan, Shuhan Wang, Xin Luo

cs.RO

2026-03-02

Putting motor temperatures and a CBF thermal reward into quadruped RL, Unitree A1 with 3 kg walks over 26 minutes; the no-thermal baseline overheats in about 7.

What problem this solves

Electric quadrupeds can handle rough ground, but their motors cook under cyclic high torque. Nameplate payload and endurance numbers are usually written for friendly cooling and moderate gaits. In the field, temperature hits a protection trip and the robot derates or stops. A robot arm can treat heat as a torque cap. A quadruped cannot. The same torque keeps balance, contact, and command tracking, so thermal safety and stability are the same control problem.

Hardware fixes add sinks, immersion, or thermal mass. Control-layer work mostly sits on fixed-base arms. This group at Huazhong University of Science and Technology wants the gait policy itself to stay off the thermal ceiling, with no extra cooling hardware.

Method

They drop a whole-body thermal network of the Unitree A1 into Isaac Gym. Each motor is a first-order thermal mass; heat is mostly winding I²R. Motors, the onboard computer, and the environment couple through thermal resistances, giving a 14-dimensional temperature state. The policy and the thermal model both run at 50 Hz, joint PD at 200 Hz. Because torque moves much faster than temperature, they feed the thermal model the RMS torque over each update window in place of current.

The actor sees all 12 motor temperatures. Training uses an asymmetric actor-critic: the critic also gets linear velocity and external force. An encoder reads the last six proprioceptive frames, estimates velocity, and emits a 16-D latent. Hybrid Internal Optimization updates the encoder; PPO updates actor and critic.

The temperature bonus does not simply fine crossing Tmax. Heat has inertia, so temperature barely falls inside one episode and the start temperature would dominate the return. They relax "stay below 60°C" into a control barrier: near the limit, the temperature derivative must stop pointing up. Penalties use temperatures clipped to 55–65°C, weight 2.0, γT = 0.35. The baseline is the same stack with that reward removed.

To make overheating show up in simulation they randomize payload 0–4 kg, CoM offset, persistent force ±30 N, friction 0.2–1.25, motor strength 0.8–1.2×, initial motor temperature near the threshold, and ambient 0–35°C. 8192 agents train for 6000 epochs, 20 s episodes, omnidirectional walking on flat ground with ±2 cm height noise and slopes up to 2%.

Results

On hardware a 3 kg pack sits on the back and the robot is commanded at 1 m/s on flat ground.

PolicySetupOutcome
No temperature reward3 kg, 1 m/sFront-left knee overheats, stop at about 7 min
Thermal-awaresameWalks until the battery dies, more than 26 min, all motors below Tmax

The abstract says over 27 minutes. Outdoor starts are not temperature-matched, so the curves are a trend, not a paired trial.

The mechanism study is in simulation with an extra 4 kg hanging on the upper-left torso, still at 1 m/s. The thermal policy uses smaller joint excursions and lower peak and RMS torques. Step frequency is higher, yet peak vertical ground force and time-averaged vertical impulse stay comparable, and horizontal components stay small. Steady gait has more pitch and a higher center of mass. The stance Jacobian's vertical entries are larger than the baseline, horizontal entries slightly smaller. In a trot, Fz dominates Fx and Fy, so the same torque buys more vertical support, heat falls, and tracking holds.

Why it matters

This is an incremental control result aimed at a real ops failure: being able to walk is not the same as being able to walk for a long time. No new motors, no liquid cooling, just temperature in the observation and a barrier-style reward on a PPO locomotion stack, and endurance moves from a minutes-scale thermal trip to the battery dying first. For inspection and payload hauling, that interruption is often worse than a missing bit of peak speed.

Reproduction is close to existing code. Thermal parameters come from their earlier lumped-parameter network, and the trainer sits near the Hybrid Internal Model locomotion line. Platforms that already expose motor temperature can try it.

Limitations

The authors say the policy keeps a conservative posture even when motors are still cool, which will hurt on steep slopes and stairs; they want a temperature-dependent mode switch later. Almost all tests are flat ground, 1 m/s, one 3 kg pack. No stairs, no hot-ambient comparison, no bake-off against hardware cooling. Torque, GRF, and Jacobian gaps live in figures; the text does not give percentages. The baseline only drops the temperature reward, with no explicit torque limit or MPC comparator. Tmax is 60°C, and how that sits against the vendor trip is not stated. The prose and Table I also disagree on the initial-temperature window (Tmax−35 versus Tmax−25).

Terms

Source

What people are saying

Related papers

All paper explainers