Does talking to an AI speaker help a toddler learn words?
My daughter asked the speaker why the moon follows the car. It answered immediately, fluently, and better than I would have. Then she stopped asking. Not just that question — she went quiet for the rest of the drive, the way you go quiet when a conversation has ended rather than when you are thinking.
I have spent eleven years building systems that answer questions well. That was the first time I watched one end a conversation I wanted my child to keep having.
The reflex is to ask whether the speaker is bad for her. I think that is the wrong question, or at least the wrong first one. The speaker did something no book and no television has ever done in that car: it took a turn. It waited for her to finish, responded to what she had actually said, and stopped. That is a conversation, structurally. The question worth asking is whether it is the kind of conversation that builds anything.
What the research says
The most useful study here is almost twenty years old and has nothing to do with AI. Patricia Kuhl’s group had nine-month-old American infants listen to Mandarin. Some heard it from a live person sitting in front of them; others heard identical recordings on a screen or over audio. The children who heard a live person learned to distinguish Mandarin phonemes about as well as infants raised in Taiwan. The children who got the recordings learned nothing measurable — their performance matched infants who had heard no Mandarin at all.
The audio was the same. The faces, in the video condition, were the same. What differed was whether the sound was coming from someone who was there.
A 2014 experiment sharpened the point. Roseberry and colleagues taught two-year-olds novel verbs in three conditions: in person, over a live video call, and from a prerecorded video. Children learned the words in person and over video chat. They did not learn them from the recording. The recording was not worse quality; it was not responsive. Video chat worked because the person on the other end paused when the child paused, looked where the child looked, and answered the question the child was actually asking.
The technical word for this is contingency, and it appears to be doing most of the work. Gilkerson and colleagues recorded whole days of home audio for 146 children and then followed them for a decade. What predicted language scores at ages nine to fourteen was not how many words the children heard. It was the number of conversational turns — how often an adult utterance was followed by a child utterance, and then by an adult response to that. Adult word count alone, once turns were accounted for, explained much less.
So the mechanism is not exposure to language. It is participation in it.
So what changes
This is where I have to be careful about what the research does and does not say, because a smart speaker is not the same thing as a prerecorded video.
A speaker is contingent in a narrow sense. It waits for you to finish. It responds to what you said rather than playing a fixed script. On the dimension that mattered in Kuhl’s and Roseberry’s experiments — responsiveness — it sits somewhere between a recording and a person, and no study I can find has tested that middle position with toddlers and vocabulary outcomes. Anyone who tells you the research shows smart speakers harm language development is overstating it. So is anyone who tells you they help.
But there is a second thing the turn-taking research implies, and it is not about the machine at all. Conversational turns are a scarce resource in a house. There are only so many moments in a day when a two-year-old is curious enough to ask something and an adult is free enough to answer. Every one of those moments that gets absorbed by a device is one that did not become a turn.
That is the part I keep thinking about. The speaker did not damage my daughter’s language. It ended a conversation — a conversation that, according to a decade-long cohort study, was one of the things most likely to matter.
The honest summary is this. We know contingent interaction with a person builds vocabulary, and we know non-contingent media does not. We do not know where a responsive machine falls on that line. What we do know is that the machine competes for the same finite supply of turns, and that the supply is the thing with the evidence behind it.
This week, try one thing
When your child asks a question you know a device could answer instantly, answer it badly on purpose. Say “I don’t know — what do you think?” and wait. The answer they get will be worse. The conversation will be longer. Based on what the turn-taking research measures, the length of the conversation is the part that shows up ten years later.
This is not a rule about screens. It is a rule about who finishes the sentence.
The founding wave is open.
Sources
- Roseberry, S., Hirsh-Pasek, K., Golinkoff, R. M. (2014). Two-year-olds' word learning from contingent and noncontingent video screen media. rct , n=36 . doi.org/10.1111/cdev.13511
- Gilkerson, J., Richards, J. A., Warren, S. F. (2018). Language experience in the second year of life and language outcomes in late childhood. cohort , n=146 . doi.org/10.1177/0956797617742725
- Kuhl, P. K. (2007). Television and children's language development. rct , n=32 . doi.org/10.1073/pnas.0705345105