Tencent Hunyuan releases T1: 122B MoE model RL-trained for long-horizon terminal tasks

Tencent-Hunyuan · hf · 2026-09-11

Tencent Hunyuan has released T1 (Terminal Agent Reinforcement Learning for Long-Horizon Tasks) on Hugging Face: a 122B Mixture-of-Experts model trained with reinforcement learning to execute long-horizon terminal tasks in a cloud sandbox.

The team reports state-of-the-art results, achieved via stable actor-critic optimization and out-of-distribution training.

Original post →

More from coding & agent

coding & agent channel →