Grok 4.5 Benchmark Scores Impacted by Test Set Leak
SkyLi0n · x · 2026-07-10
Reports indicate Grok 4.5 holds an advantage on the CursorBench benchmark because its training data accidentally included an early snapshot of the Cursor codebase. The team noted the exact impact on scores is unclear, but the data has been removed from future model training.
Related event: Grok 4.5 Benchmarks Strong but Faces Data Controversy(6 posts)→
More from Models
- Open-source labs could distill a state-of-the-art model to 32GB or 80GB VRAM, the post argues — bookwormengr · 2026-07-21
- Moonshot pauses Kimi K3 signups five days after launch as GPU demand surges — eyishazyer · 2026-07-21
- AI Diplomacy demo makes agents negotiate, ally, and betray each other — jamdac · 2026-07-21
- Newer models need a different prompting style, and old tricks can make outputs worse — emollick · 2026-07-21
- GLM-5.5 is said to arrive in 4 weeks with open weights — tanay_mehta · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21