Model comparison: Qwen, GLM, and Grok pass while ChatGPT and Claude fail
QuixiAI · x · 2026-08-28
A comparison test of several major LLMs shows that ChatGPT, Claude, and Deepseek failed, while Qwen 3.8, GLM 5.3, and Grok passed. The specific criteria of the test were not detailed in the tweet.
More from Models
- Voice Mode Test: ChatGPT and Gemini Detect Whispering, Grok Fails — deferare · 2026-08-28
- Comparison ranks GLM, Qwen, and DeepSeek Flash models — nijfranck · 2026-08-28
- Z.ai runs GLM-5.3-Flash entirely on Chinese AI chips, cutting costs by over 40% — yogthos · 2026-08-28
- Why AI Models Always Pick the Number 3 When Asked to Choose Between 1-4 — LChoshen · 2026-08-28
- Tavus launches Sparrow-2 for real-time conversational understanding — jasonkneen · 2026-08-28
- Sparse attention can replace global attention without downsides — stochasticchasm · 2026-08-28