PingPong: Hand-Curated Code-Switched Multi-Party Conversation Benchmark Accepted at EMNLP 2026
ponguru · x · 2026-09-14
The SARAL team's 'PingPong' paper is accepted at EMNLP 2026: a rare manually curated benchmark for realistic code-switched multi-party conversations, spanning 6 language combinations and 3 tasks. Evaluation shows current models still struggle with natural code-switching. Paper, data, and a Hindi podcast walkthrough are public.
More from Models
- Swift-Qwen3.8-27b, a token-efficient reasoning Qwen finetune, trends on Hugging Face — ukisai · 2026-09-14
- GPT-6 Astra hands-on: composes first, orchestrates later, and reportedly outshines Fable and Sol — paw_lean · 2026-09-14
- Researcher posts proof he both synthesized viruses and trained a 375B open-weight LLM — ethanCaballero · 2026-09-14
- Toby Ord: 10x more RLVR compute cuts tokens-to-target ~3x; gains may be math-specific — tobyordoxford · 2026-09-14
- New scaling curve has half the slope: 10,000x compute for 20%-to-80%, but bigger generational jumps — tobyordoxford · 2026-09-14
- Fudan NLP paper explains why max reasoning settings can backfire on SWE benchmarks — karminski3 · 2026-09-14