Grok 4.6 leads in legal/GDP benchmarks, lags in coding
ChrisGPT · x · 2026-08-20
Grok 4.6 has reached the 60s on an AI analysis index. Interestingly, it is not dominating coding benchmarks but rather excels in legal and GDP valuation benchmarks.
More from Models
- Sanity Check: Claude $100 Seat Equals ~$2500 in API Credits for Heavy Users — weed_cutter · 2026-08-21
- Claude accurately predicted Qwen 3.8 27B performance benchmarks — OneMoreName1 · 2026-08-21
- GLM 5.3 Scores 47.1% on SlopCodeBench, Ties with Fable 5 — corruptbytes · 2026-08-21
- Post-training causes LLMs to produce novel but impractical language — TuhinChakr · 2026-08-20
- User says Grok 4.6 now handles 100% of coding work previously done with Codex — CedricMakes · 2026-08-20
- Leaked System Prompt: Domestic Giant's Client Uses 3-Layer Memory & MCP Routing — vista8 · 2026-08-20