Kids learn language with 100,000x less data than AI - scientists don't know why
The data efficiency gap is the biggest unsolved problem in AI - and closing it could reshape everything from chatbots to minority language preservation.

Cognitive scientists and AI researchers are confronting the data efficiency gap: children master language with roughly 100,000 times less input than large language models. This gap challenges the scaling paradigm of AI and could force a rethink of how models are trained.
Kids learn language with 100,000x less data than AI - and scientists still don't know why. That's the data efficiency gap, the yawning divide between a toddler who picks up a mother tongue after hearing roughly 100 million words and a large language model that chews through 15 trillion tokens in pretraining alone. As Michael C. Frank, a cognitive scientist at Stanford, puts it: 'We still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.'
The gap is not just an academic curiosity. It's a looming constraint on the AI industry. Meta's Llama 3.1, released two years ago, was pretrained on 15 trillion tokens, and frontier models may be using 10 times that, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. But the internet's easily available data is finite - and could run dry as early as the 2030s. Kids show that learning more with less is possible, which makes understanding their secret a strategic imperative for AI labs.
The mystery is ancient. Humans have been talking for at least 100,000 years, and for all that time, the only perfect language learner was a human child. Now there are two - children and LLMs - but they learn in radically different ways. A preteen in a linguistically rich home may have heard around 100 million words; add literacy and that reaches maybe 300 million by age 20. Claude, by contrast, has seen 'the amount of language that an entire city will experience in one generation,' says Wilcox. Print out all the words used to train a modern LLM and the stack would reach past the International Space Station; a child's 100 million words would stack just 20 meters.
The question of how babies do it has divided linguists for decades. In the 1950s, MIT's Noam Chomsky argued that children are born with hardwired knowledge of grammar, countering B.F. Skinner's view that language is learned purely through conditioning. Chomsky's 'poverty of the stimulus' argument held that language is too complex and children's exposure too thin for pure statistical learning. That view dominated US linguistics and influenced early AI, which tried to code grammar rules explicitly - an approach that largely failed and contributed to the AI winter of the 1970s.
The pendulum swung back with neural networks and the transformer architecture. By 2018-2019, BERT and GPT-2 showed that massive data plus statistical learning could work for language. ChatGPT's 2022 breakout made it clear to everyone. But LLMs are 'powerful statistical learners - naïve pattern-learning machines without any of the evolved biological quirks folded into the human cortex.' They are not brains, and their data hunger is a fundamental limitation.
Closing the data efficiency gap has practical payoffs. More data-efficient models could be trained effectively on video, and could serve minority language communities that lack vast text corpora. It could also settle enduring questions about language: Are we born with a language instinct, or can language be learned purely from experience? Is our processing a biological quirk or a reflection of universal constraints? Testing hypotheses about human learning in machine models could answer both.
For executives, the implication is clear: the scaling era of AI may be hitting a wall. The next leap in capability may not come from more compute or more data, but from understanding how a child's brain achieves so much with so little. That's a research bet with potentially enormous returns - and one that could redefine the competitive landscape for AI labs, chipmakers, and anyone building on foundation models.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Science
NASA's Roman Telescope Launches Sunday, Aiming to Map Dark Energy
A new NASA observatory will probe the universe's accelerating expansion and hunt for exoplanets, with launch set for Sunday.
Nepal floods leave 1,300 missing, including dozens of Americans and Canadians
Families await word on loved ones as authorities search for more than 1,300 people after flash floods hit Nepal.
Ocean hits record 21.1°C as El Niño supercharges warming
A new daily sea-surface temperature record signals accelerating climate risk for coastal economies, insurers, and supply chains.



