childes-db 2026.1 Released: Child Language Database Grows to 24.2M Utterances with New Annotations
najoungkim · x · 2026-10-03
Michael Frank's lab released a new version of childes-db (2026.1), its versioned, reproducible interface to the CHILDES child language transcript database, accessible via R, Python, or SQL.
- More data: 24.2M utterances (+24%) and 89M tokens, including 73 corpora new to the database
- New annotation tiers: Universal Dependencies morphology (morpheme-level table) and dependency parses, Japanese romanization, speech acts, and phonology
- Versioned hosting: every release browsable, queryable, and downloadable on Redivis
- Corrected target-child identities across corpora (Day corrections)
A significant infrastructure update for child language acquisition and LLM-learning comparison research.
More from Research
- A robot demonstration can lose its most useful half-second: tracking gaps sabotage manipulation learning — Klutzy_Cap8492 · 2026-10-03
- Hinton: leading researchers now think an intelligence explosion may happen soon — geoffreyhinton · 2026-10-03
- Spotlight architecture scrutiny: attention-style memory may carry a large constant factor — teortaxesTex · 2026-10-03
- IEEE VIS 2026 Unveils Speaker Lineup, Adds GenAI and Agent Visualization Workshops — arvindsatya1 · 2026-10-03
- Spotlight memory architecture grows LLM capacity with linear cost, beats attention on long context — marcbhargava · 2026-10-03
- CortexRetrievalBench: New benchmark tests retrieval on financial docs with 140 diligence queries — Exp_Mark · 2026-10-03