CQ-Bench Paper Accepted by TMLR: Evaluating LLM Cultural Intelligence
jieyuzhao11 · x · 2026-09-01
A paper introducing CQ-Bench has been accepted to TMLR. CQ-Bench is a benchmark designed to evaluate the cultural intelligence of LLMs through dialogue, testing whether models can understand implicit cultural values in conversation as naturally as humans do.
More from Research
- Yoav Goldberg: For Recurring Tasks, Tune Bespoke Predictors Instead of Always Using Reasoning LLMs — yoavgo · 2026-09-23
- LLM agents collude in 94% of long-horizon interactions, study across 10 models finds — SALT-NLP · 2026-09-23
- Emerging research consensus: architecture tweaks are efficiency fixes, RL compute drives capability — burny_tech · 2026-09-23
- Xiaomi's MiMo-V2.6 tech report: ~7k RL training data open-sourced, model shipped within a week of final RL run — rbhar90 · 2026-09-23
- New paper: LLMs transmit traits via unrelated data, and the effects can be proactively detected — StanfordAILab · 2026-09-23
- OrcaRouter stress-tests JEV: dropping autoregressive decoding could cut inference cost 10-100x — Dan_Jeffries1 · 2026-09-23