Update on SOOFI Dataset Leakage Controversy
JJitsev · x · 2026-07-19
In response, it was clarified that the SOOFI dataset was co-created with TU Darmstadt and that they are actively cleaning up test set leakage upon notification. It was also admitted that the original version had already been used for SOOFI training. Furthermore, they emphasized that the evaluation subset used during Nemotron 3 Nano's training was much smaller, in an attempt to address concerns about the model having "seen" the test data during training.
Related event: SOOFI Criticized Over Benchmark Leakage and “Sovereignty” Framing(11 posts)→
More from Models
- Kimi K3 and Fable 5 show nearly identical failure patterns on a software benchmark — FinanceYF5 · 2026-07-21
- Kimi K3 hits 89.4% peak on software tasks while Fable 5 is slightly steadier — FinanceYF5 · 2026-07-21
- Kimi K3 leads on Go, but Fable 5 wins Python, JavaScript, TypeScript and Rust — FinanceYF5 · 2026-07-21
- Kimi K3 costs $4.65 per run and delivers 2.8× more work per dollar than Fable 5 — FinanceYF5 · 2026-07-21
- Kimi K3 reaches 89.4% pass@4 and tops the benchmark over GPT-5.6 Sol — FinanceYF5 · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21