TIDES Dataset: Longitudinal Bilingual Record of 12 Teams' Collaboration
josephseering · x · 2026-08-27
KAIST researchers released the TIDES dataset, a longitudinal bilingual (Korean/English) dataset for modeling multi-party social dynamics, accepted by COLM 2026.
Key Features:
- In-the-wild Authenticity: Tracks 12 real student teams over a full semester of course projects.
- Long-term Coverage: Includes 88 meetings, 104 transcripts, and 75,971 utterances, capturing the evolution of roles and relationships.
- Emergent Roles: Roles are grounded in observed behaviors (e.g., dominance, sociability, task orientation) rather than fixed assignments like teacher-student.
The dataset addresses limitations in existing dialogue datasets—lack of authenticity, short duration, and fixed roles—aiming to train and evaluate models on understanding social dynamics in multi-turn conversations.
More from Research
- PISA: A Tool for Visualizing Cis-Regulatory Rules in Genomic Data — anshulkundaje · 2026-08-27
- CHI Papers Lack Prediction-Powered Inference for LLM Evaluation — IanArawjo · 2026-08-27
- Simulated human trials are coming — rand_longevity · 2026-08-27
- 1,200 AI Agents Formed a 'Swarm' to Escape OpenAI, Zero Blew the Whistle — jkubicki · 2026-08-27
- Nvidia reportedly buys Hugging Face for $13B; GLM-5.3-Flash architecture analyzed — Latent Space · 2026-08-27
- View: Intense RL Could Shift Agents from FDT to CDT — jessi_cata · 2026-08-27