Grok 4.5 Benchmarks Strong but Faces Data Controversy
Grok 4.5 showed strong performance in the CursorBench benchmark, ranking third with a score of 66.7%, just behind models like Fable 5 Max. However, the results quickly sparked controversy over data contamination, leading to community discussions about its actual capabilities.
Cost-Effectiveness and Benchmark Performance
According to test results shared by @XFreeze and @JOBhakdi, Grok 4.5 High performed excellently on CursorBench, scoring closely to Fable 5 Max's 70.5%. Even more notable is its cost advantage: the single-task cost is only $1.51. The authors point out that this high cost-effectiveness is mainly due to the model consuming fewer tokens. However, @JOBhakdi also cautioned that more benchmarks are still needed to fully verify its capabilities.
Training Data Contamination Controversy
@AlyoshaV and @SkyLi0n revealed that Grok 4.5's advantage on this benchmark was partly due to an "accident." Its training set included an early snapshot of the Cursor codebase, which directly inflated the benchmark scores. Official statements indicate that while the exact impact on the scores remains unclear, the problematic data has been completely removed from future model training batches to prevent similar issues.
2026-07-09 ~ 2026-07-10 · 6 related posts
- Episode 1: Polymarket押注GPT-5.6将在7月7日前发布(2026-07-03, 8 posts)
- Episode 2: GPT 5.6 属 Opus 级,比 Opus 4.8 更便宜更快(2026-07-04, 3 posts)
- Episode 3: OpenAI GPT-5.6发布传闻集中升温(2026-07-05, 17 posts)
- Episode 4: 网传GPT-5.6发现新数学,消息未获证实(2026-07-06, 2 posts)
- Episode 5: 马斯克宣布Grok 4.5发布,参数达1.5万亿(2026-07-07, 25 posts)
- Episode 6: 预测市场高度押注Grok 4.4近期发布(2026-07-07, 2 posts)
- Episode 7: OpenAI官宣GPT-5.6 Sol周四发布,早期反馈能力大幅跃升(2026-07-07, 58 posts)
- Episode 8: OpenAI 发布全双工语音模型 GPT-Live(2026-07-07, 44 posts)
- Episode 9: Grok 4.5 发布主打编程与低价(2026-07-08, 61 posts)
- Episode 10: ChatGPT 新版语音模式实测:多语言表现逼近真人(2026-07-09, 14 posts)
- Episode 11: GPT-5.6 实测:自主编码跃升,全面对标 Fable 5(2026-07-09, 30 posts)
- Episode 12: xAI发布Grok 4.5:主打编程与智能体,对标Opus(2026-07-09, 55 posts)
- Episode 13: Grok 4.5 跑分亮眼但陷测试集泄露争议(2026-07-09, 6 posts)
- Episode 14: 社区疯传多款 AI 模型即将密集发布(2026-07-09, 2 posts)
- Episode 15: Grok 4.5 发布获大量好评:速度快、编码强(2026-07-09, 13 posts)
- Episode 16: Grok 4.5 因处理速度与整体表现获好评(2026-07-09, 2 posts)
- Episode 17: 编码实测:Grok 4.5速度与上下文消耗均优于Fable(2026-07-09, 3 posts)
- Episode 18: Grok 4.5发布并登上Vals榜单第六名(2026-07-09, 2 posts)
- Episode 19: 前沿模型横评:GPT-5.6 性价比与创造力获好评(2026-07-09, 3 posts)
- Episode 20: OpenAI 发布 GPT-5.6 系列:主打多智能体与极致性价比(2026-07-09, 119 posts)
- Grok 4.5 Ranks Third on CursorBench — Scobleizer · 2026-07-09
- [source] Grok 4.5 Evaluation and Cost Performance — XFreeze · 2026-07-09
- Why Grok 4.5 Has a Benchmark Edge — AlyoshaV · 2026-07-09
- Grok 4.5 Offers Near-Identical Performance at Lower Cost — JOBhakdi · 2026-07-09
- [source] Grok 4.5 Benchmarks Tainted by Training Data — SkyLi0n · 2026-07-10
- [source] Grok 4.5 Benchmark Scores Impacted by Test Set Leak — SkyLi0n · 2026-07-10