Harrier Sparks Controversy on MMTEB Leaderboard
antoine_chaffin · x · 2026-07-15
The post states that Harrier models achieved MMTEB top-1 in their parameter scale group and scored 62.18 / 66.35 / 69.32 on MIRACL.
However, the author argues that the current MMTEB leaderboard display does not align with user expectations for finding a good multilingual model:
- Some models are not strong overall but rank high due to using training sets from certain task datasets.
- This suggests that leaderboard results may reflect dataset memorization rather than generalization.
Subsequent replies further point out that two official MMTEB tasks themselves suffer from this issue.
Related event: MMTEB Leaderboard Faces Training Data Contamination Controversy(3 posts)→
More from Models
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22
- Google DeepMind launches Gemini 3.5 Flash Cyber for faster, cheaper code security — ralucaadapopa · 2026-07-22
- Poolside’s Laguna S 2.1 gets a two-week free run on Nous Portal — NousResearch · 2026-07-22
- Qwen3.8 Max Preview looks substantially better in a side-by-side test with Kimi K3 — curiousily_ · 2026-07-22
- Moonshot’s Kimi K3 reaches #5 on MathArena as the top open model — xeophon · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22