LLMs are too helpful: Should we include child language in training data?

Heavy_Carpenter3824 · reddit · 2026-08-29

The author found that while prototyping a child-simulating AI, current LLMs are too helpful and linguistically mature because they are trained on adult data.

The post raises a question: Should we include baby, toddler, and child speech in training datasets? Despite privacy concerns, childhood interaction is rich in discovery language and precise questions based on incomplete world models. It also correlates with behavioral shifts like the separation of imagination and reality (hallucination). TLDR: Is there important linguistic and logical structure in childhood language that models miss?

Original post →

More from AGI Musings

AGI Musings channel →