Cohere Launches Multilingual Model Human Evaluation, Seeks Contributors
Cohere_Labs · x · 2026-08-01
Cohere Labs announced that its multicultural riddle benchmark project has entered the human evaluation phase, with over 61,000 model responses already evaluated.
The project is actively recruiting native speakers for subsequent evaluations, with a specific need for contributors in Arabic dialects (Morocco and Sudan). Participants can engage in human evaluation, benchmark testing, and result analysis, with opportunities to co-author the final publication.
Related event: Cohere's Multicultural Riddle Benchmark Enters Human Evaluation(2 posts)→
More from Research
- CWoMP accepted to EMNLP 2026: Interpretable retrieval-based glossing for endangered languages — fredahshi · 2026-08-24
- SemiAnalysis Open Sources $3M AgentX Benchmark for Agentic Coding Workloads — AccBalanced · 2026-08-24
- Vinci2 Agent Outperforms GPT-5-mini in Proactive Assistance Benchmark — jiqizhixin · 2026-08-24
- OpenAI hiring for Economics of Transformative AI, MATS fellowship applications open — Astral Codex Ten · 2026-08-24
- New "Discovery Episode" Framework Measures AI Scientists by Real Research Cycles — 量子位 · 2026-08-24
- AI Claims Breakthrough on Erdős Problem Transcendence — inductionheads · 2026-08-24