Grok 4.7 flops in 100 multi-agent coding evals despite insightful solutions

teortaxesTex · x · 2026-09-23

In 100 multi-agent coding evaluations, Grok 4.7's working solutions were often more insightful than comparable frontier models, but it wasted turns and whole submissions on syntax and runtime errors and responded much slower than Grok 4.6 — overall a flop, per the tester. Third-party and unverified.

Original post →

More from Models

Models channel →