Grok 4.6 Review: Closes Gap with Qwen and Kimi, But Lags Behind Top Proprietary Models
bindureddy · x · 2026-08-14
A user has published a hands-on evaluation of Grok 4.6's latest capabilities:
- Open-Source Benchmark: Grok 4.6 scores just below top open-source models like Qwen and Kimi K3. However, it is considered more usable than K3 due to faster generation speeds.
- Capability Ceiling: The author emphasizes that despite any hype, Grok 4.6 is nowhere near the tier of top proprietary models like Claude 3.5 Sonnet (referred to as Fable) or GPT-4 (referred to as Opus).
Related event: Grok 4.6 Tested: Fast and Approaching Top-Tier Performance(4 posts)→
More from Models
- Inception Offers YC Startups 250B Free Tokens to Push Diffusion LLMs — volokuleshov · 2026-08-14
- Open Weights Model Usage on OpenRouter Drops Below 50% — maferase · 2026-08-14
- Gemini Flash 3.7 scores just below Kimi K3, but instruction-following remains weak — bindureddy · 2026-08-14
- Prediction: DeepSeek Will Cut Prices Again Once New Compute Arrives — teortaxesTex · 2026-08-14
- Google Rolls Out Gemini 3.7 Flash: Generates Playable 90s Games from a Single Prompt — googledevs · 2026-08-14
- Small Models Beat Large Ones in VLM Grounding with Tool Use — mervenoyann · 2026-08-14