Grok 4.5 Benchmarks Strong but Faces Data Controversy

Grok 4.5 showed strong performance in the CursorBench benchmark, ranking third with a score of 66.7%, just behind models like Fable 5 Max. However, the results quickly sparked controversy over data contamination, leading to community discussions about its actual capabilities.

Cost-Effectiveness and Benchmark Performance

According to test results shared by @XFreeze and @JOBhakdi, Grok 4.5 High performed excellently on CursorBench, scoring closely to Fable 5 Max's 70.5%. Even more notable is its cost advantage: the single-task cost is only $1.51. The authors point out that this high cost-effectiveness is mainly due to the model consuming fewer tokens. However, @JOBhakdi also cautioned that more benchmarks are still needed to fully verify its capabilities.

Training Data Contamination Controversy

@AlyoshaV and @SkyLi0n revealed that Grok 4.5's advantage on this benchmark was partly due to an "accident." Its training set included an early snapshot of the Cursor codebase, which directly inflated the benchmark scores. Official statements indicate that while the exact impact on the scores remains unclear, the problematic data has been completely removed from future model training batches to prevent similar issues.

2026-07-09 ~ 2026-07-10 · 6 related posts

Full story(20 episodes)→