Hy3 Beats GLM-5.1 in Blind Tests, Excels in Frontend and CI/CD

eliebakouch · x · 2026-07-06

According to the Hy3 model card, a blind test involving 270 domain experts yielded 312 valid comparisons. Hy3 achieved an average score of 2.67/4, outperforming GLM-5.1's 2.51/4. The model's advantages are most prominent in frontend development, CI/CD, and data/storage tasks. The developers explicitly stated that public leaderboards do not reflect the full picture, which is why they opted for testing based on real-world workflows.

Related event: Tencent Open-Sources Hunyuan Hy3 for Agentic Workloads(26 posts)→

Original post →

More from Models

Models channel →