Model comparison: Qwen, GLM, and Grok pass while ChatGPT and Claude fail

QuixiAI · x · 2026-08-28

A comparison test of several major LLMs shows that ChatGPT, Claude, and Deepseek failed, while Qwen 3.8, GLM 5.3, and Grok passed. The specific criteria of the test were not detailed in the tweet.

Original post →

More from Models

Models channel →