GLM 5.3 Flash vs Tencent Hy3: A Sycophancy Test Crowns Two Least Sycophantic Models
ramendik · reddit · 2026-10-06
After testing model sycophancy, the author found a clear duo of winners: GLM 5.3 Flash (which in his smoke tests is somehow less sycophantic than full GLM 5.3) and Tencent Hy3, surfaced via the lechmazur/sycophancy benchmark repo.
Early observations: Hy3 has a tighter style but tends to lose detail (less so with search enabled), while GLM 5.3 Flash is more exact but stylistically generic. Published benchmarks clearly favor GLM 5.3 Flash, though the author cautions against trusting benchmarks blindly. He's soliciting real-world opinions on both models for agentic loops, coding, general assistance, and creative writing.
More from Models
- Dev burns $700 in Opus 5.5 credits on one weekend build: 'never seen a model do that' — ivan_bezdomny · 2026-10-06
- Anthropic reviewers alerted police to a Claude chat threatening a sheriff's office, leading to an arrest — rohanpaul_ai · 2026-10-06
- flow-1: RL-trained model matches GPT-6-sol at trace debugging while 23x cheaper — kalyan_kpl · 2026-10-06
- First large-scale 3B/8B continuous diffusion LMs match pass@1 and beat pass@k vs masked dLMs — ArashVahdat · 2026-10-06
- Claude nitpicks, Codex says LGTM: what happens when AI rivals review each other's PRs — _lewtun · 2026-10-06
- OpenAI boosts default GPT-6 Astra and GPT-6.1 Sol speed ~50% to 50 tokens/sec — kimmonismus · 2026-10-06