儿童语言学习效率远超AI,原因成谜
Kids outlearn AI—and we still don’t know why
People have been talking to each other for at least 100,000 years, as best we can tell. And in all that time, there has been only one thing in the world that could learn a human language to perfect fluency: a human child.
据我们所知,人类至少已经互相交流了10万年。而在整个这段历史中,世界上只有一种东西能够以完美的流利度习得人类语言:那就是人类儿童。
Now there are two.
现在有了两种。
Four short years after the release of ChatGPT, many of us now take it for granted that we can converse naturally with our phones or computers. LLMs like Claude, DeepSeek, and OpenAI’s GPT models are fluent and flexible enough to masquerade convincingly as humans. But peek behind the computational curtain, and there’s a catch: Teaching a computer to use human language still requires an inhuman amount of data. An LLM can easily churn through a hundred thousand times more words than a person will experience in the process of mastering their mother tongue—and way more than children might hear by their first birthday, when they typically start to grab hold of language.
ChatGPT发布短短四年后,我们许多人如今已理所当然地认为,我们可以与手机或电脑进行自然的对话。像Claude、DeepSeek以及OpenAI的GPT模型这样的大语言模型(LLMs)足够流利且灵活,能够令人信服地伪装成人类。但一旦窥探计算幕后的真相,就会发现一个缺陷:教计算机使用人类语言仍然需要非人般庞大的数据量。一个LLM可以轻松处理比一个人掌握母语过程中所经历的词汇量大十万倍的文字——这远远超过儿童在一岁左右开始掌握语言时所能听到的词汇量。
“The progress recently has been amazing,” Michael C. Frank, a cognitive scientist at Stanford University, says of LLMs. “But we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year.”
斯坦福大学认知科学家迈克尔·C·弗兰克(Michael C. Frank)在谈及LLMs时表示:“最近的进展令人惊叹。但我们仍然必须砍伐一片森林并搜集全人类知识的总和,才能重现这一在我们客厅里一年内就会发生的里程碑事件。”
This yawning divide between children and machines is called the data efficiency gap. And it raises a tantalizing question for cognitive scientists and a challenge for the architects of AI models: How is it that kids can still outperform the most linguistically sophisticated machines ever built?
儿童与机器之间的巨大鸿沟被称为“数据效率差距”(data efficiency gap)。这对认知科学家提出了一个诱人的问题,也对AI模型的架构师构成了挑战:为什么孩子们依然能超越有史以来最精通语言的机器?
Finding answers has stakes for both AI research and cognitive science. For the past decade, language models have mostly gotten better by getting bigger. Meta’s open-weight LLM Llama 3.1, released two years ago, chewed through 15 trillion tokens (word-like chunks of language) in pretraining—the main step of training a model that happens before it is fine-tuned for a specific task, like being a chatbot. Frontier models could be pretraining on 10 times more data, says Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University. But there’s only so much internet to train on, and eventually—perhaps as early as the 2030s—the well of easily available data could run dry.
寻找答案对AI研究和认知科学都具有重要意义。在过去十年里,语言模型主要通过变得更大来提升性能。两年前发布的Meta开源权重LLM Llama 3.1,在预训练阶段(即在对特定任务如充当聊天机器人进行微调之前进行的模型主要训练步骤)处理了15万亿个token(类似单词的语言单元)。乔治城大学的认知科学家兼语言学家伊桑·戈特利布·威尔科克斯(Ethan Gotlieb Wilcox)表示,前沿模型可能在预训练中使用多达10倍的数据量。但可供训练的网络资源是有限的,最终——也许早在2030年代——容易获取的数据源泉可能会枯竭。
Kids show that it could be possible to learn more with less. Far less. A preteen raised in a linguistically rich home may have heard something in the vicinity of 100 million words. Add literacy to the mix and you can boost that word count to maybe 300 million words by age 20.
孩子们表明,用更少的资源学习更多是可能的。少得多。在一个语言环境丰富的家庭中长大的青少年,可能听到了大约1亿个单词。如果再加上读写能力,到20岁时,这个词汇量可能增加到3亿个左右。
The difference in scale is something that can only really be gestured at in analogy. “Claude has seen the amount of language that an entire city will experience in one generation,” says Wilcox. If you were to print out on paper all the words used to train a modern LLM, you could make a stack that would reach past the International Space Station. The human preteen’s 100 million words, meanwhile, would stack up just 20 meters. And we can make do with far less than that.
规模上的差异只能通过类比来大致描述。Wilcox 说:‘Claude 所接触到的语言量,相当于一个城市整整一代人所经历的语言总量。’如果你把训练现代大型语言模型(LLM)所用的所有单词打印在纸上,堆起来的高度将超过国际空间站。相比之下,人类青春期前儿童接触的约 1 亿个单词,堆起来只有 20 米高。而且,我们其实只需要远少于这个数量就能学会语言。
By reverse-engineering the way kids learn, scientists hope to be able to create more data-efficient AI models, which could be useful for everything from training AI effectively on video to creating chatbots that serve minority language communities. Testing hypotheses about human learning in machine models could also settle enduring questions about language and children’s developing minds. Are we born with a language instinct, or would it be possible, even in principle, for a child to learn language purely from experience? Is the way we process language a quirk of our biology, or might at least some of it reflect universal constraints on how languages can be used and learned?
通过逆向工程儿童的学习方式,科学家希望创造出数据效率更高的 AI 模型,这些模型可用于从有效利用视频训练 AI,到为少数语言群体创建聊天机器人等方方面面。在机器模型中测试关于人类学习的假设,也可能解决有关语言和儿童心智发展的长期争议。我们是天生具有语言本能吗?或者,原则上孩子是否可能仅凭经验学习语言?我们处理语言的方式是生物学的偶然特性,还是至少其中一部分反映了语言使用和学习的普遍约束?
The essential elements
基本要素
Most of us realize language is hard only when we try to learn a new one after childhood. The past perfect tense, rolled rs and nasal vowels, the genitive case, phrasal verbs, grammatically masculine tables and feminine spoons—many are the instruments of linguistic torment for the adult language learner. It’s typically effortless to learn our mother tongues, however. Toddlers usually start producing grammatically correct sentences after hearing something like 10 million words, or 30 million on the high end.
我们大多数人只有在童年之后尝试学习一门新语言时,才意识到语言很难。过去完成时、卷舌音和鼻化元音、属格、短语动词、语法上阳性的桌子和阴性的勺子——对于成年语言学习者来说,这些都是语言折磨的工具。然而,学习母语通常是毫不费力的。幼儿通常在听到大约 1000 万个单词后开始说出语法正确的句子,最多可达 3000 万个。
“It’s just totally miraculous,” says Frank. “If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.”
Frank 说:‘这简直完全不可思议。’‘如果你用 3000 万个单词训练 GPT-2,你得到的是一个胡言乱语生成器;而不是一个孩子。’
Exactly how babies pull this off is a mystery. Researchers know a lot about what kids learn and how they use language at different stages in development, but there’s still a lot we don’t know. Perhaps the most enduring question is why babies can learn language at all. The syntax of human language—the rules for combining words into sentences—includes recursive, nested structures that allow us to express virtually infinite ideas with a finite lexicon of words and pieces of words. This seems like something that should be a problem for babies. They only splash about in the shallows of a fathomless ocean of language. And yet, somehow, that’s enough. From a drop, they infer the depths.
婴儿究竟是如何做到这一点的,至今仍是个谜。研究人员对儿童在不同发展阶段学到了什么以及如何使用语言已经有了很多了解,但我们仍然有很多未知之处。或许最持久的问题是,为什么婴儿能够学习语言。人类语言的句法——将单词组合成句子的规则——包括递归的嵌套结构,使我们能够用有限的词汇和词素表达近乎无限的想法。这看起来应该是婴儿难以应对的问题。他们只是在语言这片深不可测的海洋的浅水区扑腾。然而,不知为何,这就足够了。从一滴水中,他们推断出了深海。
One solution, put forward in the 1950s by the MIT linguist Noam Chomsky, is that babies are born with hardwired knowledge of grammar. Chomsky was reacting to a rival view, championed by the psychologist B.F. Skinner, that language acquisition is entirely environmental. Skinner thought language was learned through conditioning and reinforcement, the way a dog figures out how to sit or shake for treats. Chomsky countered by citing the “poverty of the stimulus”—the idea that language, especially syntax, is too complex and children’s exposure to it too “impoverished” for them to learn entirely from experience. “His signature argument was, essentially, that language cannot be learned on the basis purely of statistics,” says Richard Futrell, a linguist and cognitive scientist at the University of California, Irvine. Instead, Chomsky posited that language is based on a set of logical rules and argued that children needed innate knowledge of those rules to deduce the grammar of their language from scraps of speech.
20世纪50年代,麻省理工学院的语言学家诺姆·乔姆斯基(Noam Chomsky)提出了一种解决方案,即婴儿天生就具备语法方面的硬连线知识。乔姆斯基是在回应一种由心理学家B.F.斯金纳(B.F. Skinner)推崇的对立观点,该观点认为语言习得完全是环境决定的。斯金纳认为,语言是通过条件反射和强化来学习的,就像狗学会坐下或握手以换取零食一样。乔姆斯基则引用了“刺激贫乏论”作为反驳——即语言,尤其是句法,过于复杂,而儿童接触到的语言输入过于“贫乏”,不足以让他们仅凭经验完全学会。加州大学尔湾分校的语言学家兼认知科学家理查德·富特雷尔(Richard Futrell)表示:“他的标志性论点本质上是,语言不能仅仅基于统计数据进行学习。”相反,乔姆斯基假设语言基于一套逻辑规则,并认为儿童需要先天具备这些规则的知识,才能从零碎的话语中推导出其语言的语法。
“It’s just totally miraculous … If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.”
“这简直太神奇了……如果你用3000万词训练GPT-2,你得到的是一个胡言乱语生成器;而你得不到一个孩子。”
Michael C. Frank, cognitive scientist, Stanford University
迈克尔·C·弗兰克(Michael C. Frank),斯坦福大学认知科学家
The Chomskyan view of language dominated linguistics in the US for decades under the moniker of generative grammar. And it was a major influence on computer science in the 1950s and ’60s, when AI was enjoying its first boom time and the lines between linguistics and natural-language processing dissolved in a flood of military funding; the Pentagon wanted computers that could understand English and translate Russian.
在“生成语法”这一名称下,乔姆斯基式的语言观在美国语言学界主导了几十年。它在20世纪50年代和60年代的计算机科学中也产生了重大影响,当时人工智能正经历第一次繁荣期,语言学与自然语言处理之间的界限在大量军事资金的涌入中逐渐消融;五角大楼希望计算机能够理解英语并翻译俄语。
Despite early successes of simple neural networks, which learn to recognize and reproduce statistical patterns, AI researchers in the United States largely adopted a rule-based framework influenced by Chomsky’s theories. They tried to teach language to computers by explicitly coding the rules into programs—think less immersion experience, more grammar class. This approach, part of a broader trend called symbolic AI, prevailed for decades. It also largely failed to produce models actually capable of handling human language at scale. Interest in natural-language processing chilled in the “AI winter” that began in the 1970s.
尽管简单神经网络在识别和复现统计模式方面取得了早期成功,但美国的人工智能研究人员在很大程度上采用了受乔姆斯基理论影响的基于规则的框架。他们试图通过将规则显式编码到程序中来教会计算机语言——这更像是语法课,而非沉浸式体验。这种方法作为更广泛的所谓符号人工智能趋势的一部分,主导了数十年。然而,它也未能大规模产生真正能够处理人类语言的模型。自20世纪70年代开始的“AI寒冬”期间,对自然语言处理的兴趣逐渐冷却。
In the aftermath, neural networks started to make a comeback. But it wasn’t until the 2010s, when computer hardware was getting cheap and capable and the internet was getting big, that their performance began turning heads. By 2018 and 2019, the models BERT and GPT-2, which were built on a new architecture—the transformer—and trained on billions of tokens, made it clear to insiders that learning from a massive glut of data could work for language. In 2022, with the breakout success of OpenAI’s chatbot ChatGPT, it was clear to everyone.
在此之后,神经网络开始复兴。但直到2010年代,随着计算机硬件变得廉价且强大、互联网规模不断扩大,它们的性能才开始引人注目。到了2018年和2019年,基于全新架构——Transformer——并在数十亿个词元上训练而成的BERT和GPT-2模型,让业内人士清楚地认识到,从海量数据中学习的方法在语言处理上是可行的。2022年,随着OpenAI的聊天机器人ChatGPT取得突破性成功,这一点对所有人都显而易见。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力