T1: RL-trained 122B MoE terminal agent lifts Terminal-Bench 2.1 from 43.8% to 64.0%

heghbalz · x · 2026-09-11

Researchers introduce T1, a 122B-total/10B-active MoE model trained with reinforcement learning to operate a real shell in a cloud sandbox for 300+ tool-call turns per task.

Key results

Training recipe

Paper, project page, and data repo are public. The authors note this echoes DeepSeek-V4.1's finding that better data and environment pipelines unlock large gains with established RL methods.

Related event: Tencent Hunyuan Releases T1, an RL-Trained Terminal Agent(2 posts)→

Original post →

More from coding & agent

coding & agent channel →