Grok 4.7 lands at #21 on HLE as GLM-5.3 and Qwen 3.8 top the charts
himanshustwts · x · 2026-09-22
- Grok 4.7 ranks only #21 on Humanity's Last Exam, per the post.
- GLM-5.3, Qwen 3.8, Muse 1.3, and DeepSeek v4.1 lead the chart, and the same models top TB 4.0.
- The post highlights Chinese/open models' clear lead on reasoning benchmarks.
Related event: Grok 4.7 Ranks Only 21st on HLE Leaderboard(2 posts)→
More from Models
- Paradigm launches Limite 1B - Violetto, a model for high-frequency mathematical intelligence — tensorqt · 2026-09-22
- Reliquary-4B: A 4B math & code model trained via decentralized RL with community rollouts — const_reborn · 2026-09-22
- Users say they can't trick Jev into hallucinating — BLUECOW009 · 2026-09-22
- Measured trade-offs of three REAP-pruned Qwen3.8-Flash-Next MLX builds on Apple Silicon — MensaProdigy · 2026-09-22
- Dev claims further-optimized DeepSeek V4 NVFP4 uses 190GB of 192GB VRAM — HankYeomans · 2026-09-22
- OpenAI researcher Will Depue on why voice models still lack true realtime chat — willdepue · 2026-09-22