Grok 4.7 flops in 100 multi-agent coding evals despite insightful solutions
teortaxesTex · x · 2026-09-23
In 100 multi-agent coding evaluations, Grok 4.7's working solutions were often more insightful than comparable frontier models, but it wasted turns and whole submissions on syntax and runtime errors and responded much slower than Grok 4.6 — overall a flop, per the tester. Third-party and unverified.
More from Models
- Mathematicians' open letter demands OpenAI release proofs of 100+ solved problems — panickssery · 2026-09-23
- Steering-vector loom: turning n completions into vectors to steer model output — repligate · 2026-09-23
- Model reviewers juggle 2-5 subscriptions per provider — corporate devs say that's not reality — emaayan · 2026-09-23
- Is DeepSeek's RL actually good? Question raised over V4.1 vs V2.6 scale gap — teortaxesTex · 2026-09-23
- GPT-6 Sol and Luna appear in Codex CLI but are missing from the VS Code extension — nambirseyr · 2026-09-23
- computer-10: a character model built on Llama 3.1 70B base with no assistant training data — liminal_bardo · 2026-09-23