Tencent’s Hy3 ranks 25th in Agent Arena after 8,000+ live agent sessions
arena · x · 2026-07-22
Tencent’s Hy3 lands at #25 in Agent Arena with a mixed scorecard
Agent Arena’s latest leaderboard puts Hy3 at #25 overall after more than 8,000 live agentic sessions.
- Overall rank: 25
- Confirmed Success: 27th, with -4.6% net improvement
- Praise vs Complaint: 20th, -2.9%
- Steerability: 30th, -7.1%
- Bash Recovery: 25th, +2.6%
- Tool Hallucination: 25th, +0.9%
The post is essentially a snapshot of where Hy3 stands on real agent tasks: mediocre overall, weaker on success and steerability, but somewhat better at recovering from shell errors.
More from coding & agent
- Harnesses widen what LMs can do, but may not improve compositional generalization — a1zhang · 2026-07-23
- A workflow for sharing agentic coding artifacts with teammates and pushing them to Git — rseroter · 2026-07-23
- An open-source /no-ai-slop skill removes 20+ common AI writing patterns — petergyang · 2026-07-23
- Open-source /no-ai-slop skill targets more than 20 common AI writing patterns — robleclerc · 2026-07-23
- LLMs are erasing framework-first identity labels among programmers — generativist · 2026-07-23
- Rork says Max App can build and preview iOS apps directly on iPhone — rudrank · 2026-07-23