SOOFI’s update still doesn’t fix the benchmark contamination problem
JJitsev · x · 2026-07-26
The author says the updated SOOFI report still fails to address the main problems.
- He argues the report keeps misleading the public while sidestepping the core issues.
- The update adds side fixes, but does not remove the underlying evaluation contamination.
- The thread criticizes the repeated “frontier-level” framing as overblown and unsupported.
More from Models
- Opus-5 is getting attention for its unusual vocabulary choices — adonis_singh · 2026-07-26
- Google’s Gemma team asks what capabilities people want in the next models — osanseviero · 2026-07-26
- Alibaba answers with Qwen3.8 as Kimi K3 and Chinese models keep closing the gap — emmanuelvivier · 2026-07-26
- Moonshot AI launches open-weight Kimi K3 and claims strong results against top US models — emmanuelvivier · 2026-07-26
- Silicon Valley is split on Chinese open-weight models now rivaling top U.S. systems — emmanuelvivier · 2026-07-26
- Kimi-K3 claims a perfect 6/6 on IMO 2026 Lean 4 proofs — songhan_mit · 2026-07-26