Cohere benchmarked North Small Translate on WMT26, released after training, to avoid contamination
cohere · x · 2026-10-04
In part 5/6 of its North Small Translate thread, Cohere explains its benchmarking approach: the team used WMT26 benchmarks, which were released after the model was created, ensuring the model could not have been trained on the test set.
This is a direct answer to common benchmark-contamination concerns around open translation models.
Related event: Cohere Open-Sources North Small Translate Model(3 posts)→
More from Models
- Antigravity Rolls Out Claude Opus 5.5 and Sonnet 5.5 to All Paid Users — firstadopter · 2026-10-04
- RL environments are data: they push the frontier until they saturate — Shahules786 · 2026-10-04
- GLM-powered chess bot escapes losing position and checkmates human opponent — MikePFrank · 2026-10-04
- mythos 5 called one of the most attractive language models — repligate · 2026-10-04
- Apodex 1.1 Mini goes live on Novita AI with 14-day free trial — SimonShaoleiDu · 2026-10-04
- Cloudflare's new vision model clef appears on Hugging Face as devs urge Perplexity to add it — ostrisai · 2026-10-04