Cohere benchmarked North Small Translate on WMT26, released after training, to avoid contamination

cohere · x · 2026-10-04

In part 5/6 of its North Small Translate thread, Cohere explains its benchmarking approach: the team used WMT26 benchmarks, which were released after the model was created, ensuring the model could not have been trained on the test set.

This is a direct answer to common benchmark-contamination concerns around open translation models.

Related event: Cohere Open-Sources North Small Translate Model(3 posts)→

Original post →

More from Models

Models channel →