Why Grok 4.5 Has a Benchmark Edge
AlyoshaV · reddit · 2026-07-09
A note on the CursorBench page explains that Grok 4.5's advantage on this benchmark is partly due to its training set accidentally including an early snapshot of the Cursor codebase. The post highlights this evaluation note, emphasizing that benchmark results can be compromised by training data contamination.
Related event: Grok 4.5 Benchmarks Strong but Faces Data Controversy(6 posts)→
More from Models
- Open-source labs could distill a state-of-the-art model to 32GB or 80GB VRAM, the post argues — bookwormengr · 2026-07-21
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21