PingPong benchmark at EMNLP 2026: 6 language pairs show LLMs still struggle with code-switching
ponguru · x · 2026-09-14
The PingPong paper has been accepted to EMNLP 2026. The authors describe it as a rare, manually curated benchmark for realistic code-switched multi-party conversations, spanning 6 language combinations and 3 tasks. Their finding: current models still struggle with natural code-switching.
More from Research
- Diversity-Aware Skill Routing Uses DPP to Cut Redundant LLM Agent Skill Picks — Wang Wei · 2026-09-14
- Contextual Bandit Algorithms Route Prompts to LLM Experts with Sublinear Regret — Wang Wei · 2026-09-14
- Heidelberg Laureate Forum launches daily podcast, first episode features Fintzen and Hoefler — thoefler · 2026-09-14
- WebMCP browser tools cut tokens 52% on one task but increase them on another, DeepDeck experiment finds — j032 · 2026-09-14
- AI reasoning isn't just step-by-step thinking: a thread breaks down 7 reasoning types — goyalshaliniuk · 2026-09-14
- Nautilus turns one prompt into plug-and-play robot learning workflows, as researchers question the GPT-6 hype — GeorgiaChal · 2026-09-14