Tencent Hunyuan releases T1: 122B MoE model RL-trained for long-horizon terminal tasks
Tencent-Hunyuan · hf · 2026-09-11
Tencent Hunyuan has released T1 (Terminal Agent Reinforcement Learning for Long-Horizon Tasks) on Hugging Face: a 122B Mixture-of-Experts model trained with reinforcement learning to execute long-horizon terminal tasks in a cloud sandbox.
The team reports state-of-the-art results, achieved via stable actor-critic optimization and out-of-distribution training.
More from coding & agent
- Reddit agent builders weigh tooling for conversation storage and agent observability — Srinidhi_Murali · 2026-09-11
- Utopia, an Open-Source Enterprise World Model, Hits 6.8k Stars on GitHub — adnan_hashmi · 2026-09-11
- OpenAI employee built a feedback app with ~750k data sources to guide product decisions — simpsoka · 2026-09-11
- Dev builds browser RTS game with ChatGPT Astra + Blender MCP in just 40 prompts — TheMoonMidas · 2026-09-11
- Agent builds 800 pages of nested proxies just to sign up for PyPI — voooooogel · 2026-09-11
- BAML: one open-source API for all LLM capabilities across six languages — dosco · 2026-09-11