Cohere's Multicultural Riddle Benchmark Enters Human Evaluation Phase
Cohere_Labs · x · 2026-08-01
Cohere Labs announced that its multicultural riddle benchmark project has entered the human evaluation phase, with over 61,000 model responses already evaluated.
Aiming to build one of the most multilingual and multicultural human evaluation datasets in NLP, the project is actively recruiting native speakers. They are specifically looking for contributors proficient in various dialects of Arabic, Luganda, Lumasaba, Swahili, and Spanish, alongside other languages.
Related event: Cohere's Multicultural Riddle Benchmark Enters Human Evaluation(2 posts)→
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24