Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
Ahmet Tuğrul Bayrak, Fatma Nur Korkmaz, Bekir Berker Türker, Mustafa Sertaç Türkel, Alper Kaplan
cs.CL, cs.AI
2026-08-23
A 4.23-hour unscripted Turkish dyadic corpus plus a genetic AND-OR rule reaches F1 57.0, only 7 points above always predicting a turn change.
Spoken agents still treat a pause as the end of a turn. Pause length varies wildly across speakers, so cutting in early creates overlap and waiting too long feels frozen. Turkish has almost no naturalistic corpus with primary turn-taking labels; existing resources lean toward sentiment or synthetic dialogue. Ata Technology Platforms recorded 11 unscripted dyadic conversations as Real-TurnTurk and evolved interpretable rules over multimodal features with a genetic algorithm, aiming at detectors cheap enough for real-time use.
Turn yield is not one cue. A finished question, a falling pitch, or a gaze shift can each hand the floor. The model has to represent several routes, not a single threshold.
The corpus is 4.23 hours of synchronized 1080p 16 fps video, per-speaker 48 kHz audio in most sessions, and millisecond-aligned transcripts. Speech fills 222.1 minutes; non-speech is 12.4%. Turns are labeled semi-automatically: speaker change, next segment at least 0.5 s, not a filler, text longer than three characters. That yields 1,750 positives and 3,500 duration-stratified negatives, with a ±1.0 s buffer around positives. Overlap appears in 81.4% of labeled transitions, so silence is a weak cue.
Each candidate uses only the current speaker’s previous 2 s and 28 features: 9 visual (MediaPipe Face Mesh plus Py-Feat), 12 acoustic (openSMILE eGeMAPSv02), 7 linguistic (word duration, syllables, filler, question, completeness). Listener gaze and nods are unused.
The task is binary classification. A chromosome encodes 3 to 7 conditions, each with a feature, threshold, comparator, and AND/OR. Population 500, at most 900 generations, training-fold F1 as fitness. A hit within 0.5 s of an annotated change counts as correct. Five-fold CV splits by conversation. Two hub speakers appear in every fold, so a speaker-disjoint estimate is impossible and the reported numbers are an upper bound on unseen-speaker generalization.
The evolved rule averages precision 46.1, recall 74.6, F1 57.0 on held-out conversations.
| Method | Precision | Recall | F1 |
| Always predict change | 33.3 | 100.0 | 50.0 |
| Silence 1.0 s | 10.5 | 40.3 | 16.7 |
| Silence 2.0 s | 20.4 | 35.0 | 25.8 |
| Silence 3.0 s | 22.2 | 10.7 | 14.4 |
| GA rule | 46.1 | 74.6 | 57.0 |
That is 7 F1 points above always-positive. The authors decline a significance claim on 11 conversations. Silence baselines all fall below always-positive because pauses are rare and overlap is common. The full-data readable rule has three pathways: word duration ≥0.60 s and energy rate ≥1.00; gaze change ≥0.35 and mean F0 ≥120 Hz; filler plus word duration ≥0.80 s. Five of 28 features survive; vision contributes one clause.
For Turkish or other low-resource duplex agents, the corpus is the asset: separated channels attribute overlap, and backchannels are kept out of positives. The rule is meant to be inspectable in a live stack, not a state-of-the-art classifier. F1 57 clears the always-positive line and is still far from a shippable cut-in detector.
Speaker imbalance is severe: user 01 and user 03 account for 60.8% of words and 60.1% of talk time, and every conversation contains one of them, so speaker leakage is baked in. There is no Random Forest, XGBoost, or Transformer comparison, so the cost of interpretability is unknown. Listener visuals are unused. The 1:2 negative ratio was fixed a priori. A subset sits on Hugging Face; full public release is unspecified. The authors themselves call the 7-point margin an upper bound and plan a balanced, non-hub recording design.