AI并非更会思考,而是拥有更大的符号工作记忆
AI Isn't Outthinking Mathematicians. It's Out-Remembering Them
AI Isn’t Outthinking Mathematicians. It’s Out-Remembering Them.
AI 并非比数学家更会思考,而是比他们更擅长记忆。
The key advantage may not be superior reasoning, but a virtually unlimited symbolic working memory.
关键优势可能不在于更卓越的推理能力,而在于几乎无限的符号工作记忆。
Davide Piffer
达维德·皮费尔
Aug 04, 2026
2026年8月4日
40
40
2
2
12
12
Share
分享
At the 1952 dedication of the Institute for Advanced Study computer. AI may be less like an electronic Einstein than a machine-amplified von Neumann: immense speed, breadth and symbolic memory.
在1952年高等研究院计算机的落成典礼上,AI 或许不像电子版爱因斯坦,而更像机器增强版的冯·诺依曼:拥有极快的速度、广博的知识和强大的符号记忆。
When an AI system solves a difficult mathematical problem, the usual explanation is that it has become more intelligent.
当 AI 系统解决一个困难的数学问题时,通常的解释是它变得更聪明了。
Perhaps it has absorbed millions of mathematical examples. Perhaps reinforcement learning has taught it better reasoning strategies. Perhaps it is beginning to develop something resembling genuine mathematical intuition.
也许它吸收了数百万个数学示例。也许强化学习教会了它更好的推理策略。也许它开始发展出某种类似真正数学直觉的东西。
All of these explanations may contain some truth. But they overlook a simpler possibility:
所有这些解释可能都包含一些真理,但它们忽略了一个更简单的可能性:
AI has access to a vastly larger working memory than the human brain.
AI 拥有比人脑大得多的工作记忆。
Or, more precisely, it has access to an enormous external symbolic workspace that performs many of the functions that working memory performs in humans.
或者更准确地说,它拥有一个巨大的外部符号工作空间,执行着人类工作记忆所执行的许多功能。
This difference may be especially important in mathematics.
这种差异在数学领域可能尤为重要。
A human mathematician can hold only a small number of unfamiliar elements in mind simultaneously. An AI model can keep the entire problem statement, hundreds of intermediate equations, several abandoned approaches, definitions, constraints and earlier conclusions inside its context window.
人类数学家只能同时在脑海中记住少量不熟悉的元素。而 AI 模型可以在其上下文窗口内保留整个问题陈述、数百个中间方程、几种被放弃的方法、定义、约束和早期结论。
We normally interpret the resulting performance as evidence of superior reasoning. But some of it may instead reflect the removal of one of the most important biological limits on human reasoning: our extremely restricted working-memory capacity.
我们通常将由此产生的表现解释为卓越推理能力的证据。但其中一部分可能反而反映了人类推理中最重要的生物限制之一的消除:我们极其有限的工作记忆容量。
Mathematics is constrained by memory
数学受限于记忆
Working memory is the mental system that allows us to hold and manipulate information over short periods.
工作记忆是让我们在短时间内保持和操作信息的心理系统。
When solving an equation, you must remember what each variable represents, which operations have already been performed and what the current goal is. During a proof, you may need to keep track of assumptions, intermediate lemmas, exceptions and multiple possible cases.
在解方程时,你必须记住每个变量代表什么,已经执行了哪些运算,以及当前的目标是什么。在证明过程中,你可能需要跟踪假设、中间引理、例外情况和多种可能的情况。
Human working memory is remarkably limited.
人类的工作记忆非常有限。
Its exact capacity depends on the task and on how information is organized, but the general limitation is obvious from everyday experience. Try multiplying two three-digit numbers in your head. The underlying operations are simple. The difficulty comes largely from having to preserve partial results while performing additional calculations.
其确切容量取决于任务和信息组织方式,但一般限制从日常经验中显而易见。试着在脑海中计算两个三位数的乘法。底层运算很简单,困难主要来自于在执行额外计算时必须保留部分结果。
Writing the numbers down transforms the problem.
把数字写下来会改变问题。
Paper does not make you more intelligent. It expands your effective working memory.
纸张并不会让你变得更聪明。它扩展了你的有效工作记忆。
The same principle applies at higher levels of mathematics. A mathematician uses notation, scratch paper, diagrams and previously written lemmas not merely to communicate the solution, but to make the reasoning cognitively possible.
同样的原则也适用于更高层次的数学。数学家使用符号、草稿纸、图表和先前写下的引理,不仅仅是为了交流解决方案,而是为了使推理在认知上成为可能。
Experts compensate through “chunking.” A novice sees a long sequence of symbols. An expert recognizes a familiar structure and treats it as a single conceptual object. This allows far more information to fit inside the same biological working-memory limit.
专家通过“组块化”来弥补。新手看到一长串符号,而专家则识别出熟悉的结构,并将其视为一个单一的概念对象。这使得更多信息能够适应相同的生物工作记忆限制。
But chunking does not eliminate the limit. It merely compresses the information.
但组块化并没有消除限制,它只是压缩了信息。
An AI model faces a very different constraint.
人工智能模型面临一个非常不同的约束。
Working memory predicts mathematical performance beyond IQ
工作记忆预测数学表现,超越智商。
The importance of working memory for mathematics is not merely theoretical. It is visible in the differences between human beings.
工作记忆对数学的重要性不仅仅是理论上的。它在人与人之间的差异中可见。
Working memory is strongly related to general intelligence, which raises an obvious question: does it independently predict mathematical performance, or is it merely another imperfect measure of IQ?
工作记忆与一般智力密切相关,这引发了一个明显的问题:它是否独立预测数学表现,还是仅仅是智商的一个不完美的衡量标准?
Several studies suggest that it contributes something beyond conventional intelligence measures. Alloway and Passolunghi (2011), for example, examined working memory, verbal ability and mathematical skills in children. They found that working-memory measures made a distinct contribution to mathematical performance rather than simply reproducing the association between mathematics and general verbal ability.
几项研究表明,它贡献了超越传统智力测量的东西。例如,Alloway 和 Passolunghi(2011)研究了儿童的工作记忆、言语能力和数学技能。他们发现,工作记忆测量对数学表现做出了独特的贡献,而不仅仅是重现了数学与一般言语能力之间的关联。
In a separate six-year longitudinal study, Alloway and Alloway (2010) measured children at age five and then examined their academic achievement six years later. Early working-memory performance predicted later literacy and numeracy even after IQ was included in the analysis. Indeed, working memory was a stronger predictor of the later academic outcomes than the IQ measure used in the study.
在一项单独的为期六年的纵向研究中,Alloway 和 Alloway(2010)在儿童五岁时进行了测量,然后在六年后检查了他们的学业成就。即使在分析中包含了智商,早期的工作记忆表现也能预测后来的读写和计算能力。事实上,工作记忆对后来学业成果的预测力比研究中使用的智商测量更强。
Blankenship and colleagues (2015) similarly reported that working memory explained unique variation in mathematical fluency and calculation after statistically controlling for IQ and age. A large meta-analysis by Friso-van den Bos and colleagues (2013) also found a consistent relationship between working memory and mathematics across primary-school studies, although the strength of the relationship varied according to the type of working-memory and mathematical task being measured.
Blankenship 及其同事(2015)同样报告,在统计控制智商和年龄后,工作记忆解释了数学流畅性和计算中的独特变异。Friso-van den Bos 及其同事(2013)进行的一项大型元分析也发现,在小学研究中,工作记忆与数学之间存在一致的关系,尽管关系的强度因所测量的工作记忆和数学任务的类型而异。
These findings should not be exaggerated. Working memory and intelligence overlap substantially, and statistical control cannot perfectly isolate them as independent psychological mechanisms. Nor does the evidence imply that commercially training working memory will necessarily produce large improvements in intelligence or mathematics.
这些发现不应被夸大。工作记忆与智力有大量重叠,统计控制无法完美地将它们作为独立的心理机制区分开来。证据也不意味着商业训练工作记忆必然能大幅提升智力或数学能力。
The narrower conclusion is nevertheless important: among children with similar measured intelligence, differences in the ability to hold, update and manipulate information still predict differences in mathematical performance.
然而,更狭窄的结论仍然重要:在测量智力相似的儿童中,保持、更新和操作信息的能力差异仍能预测数学表现的差异。
This provides a crucial clue for understanding AI. If human mathematical performance is partly capped by a working-memory bottleneck, then giving a machine an enormous symbolic workspace changes the nature of the contest. The machine may appear more mathematically intelligent partly because it is much less constrained by a cognitive limitation that suppresses human performance.
这为理解人工智能提供了关键线索。如果人类数学表现部分受限于工作记忆瓶颈,那么给机器一个巨大的符号工作空间改变了竞赛的性质。机器可能显得更具数学智能,部分原因是它较少受到抑制人类表现的认知限制。
The context window is a gigantic notebook
上下文窗口是一个巨大的笔记本
A modern language model can process an enormous sequence of tokens at once. This sequence may include the original question, definitions, examples, intermediate calculations and the model’s own earlier reasoning.
现代语言模型可以一次处理大量的token序列。这个序列可能包括原始问题、定义、示例、中间计算以及模型自身的早期推理。
The context window is not identical to human working memory. It is better understood as a gigantic external notebook combined with an imperfect system for searching and using what has been written in it.
上下文窗口并不等同于人类的工作记忆。它更应被理解为一个巨大的外部笔记本,加上一个不完美的搜索和使用其中内容的系统。
This distinction matters.
这种区别很重要。
Humans possess a form of active internal memory. We can silently choose a number, hold it in mind, transform it and replace it with a new value without saying or writing anything.
人类拥有一种主动的内部记忆形式。我们可以默默选择一个数字,在脑海中保持它,转换它,并用新值替换它,而无需说出或写出任何内容。
Standard language models are much weaker at maintaining this kind of private, continuously updated mental state. Their most stable form of memory is usually the sequence of tokens that has already been generated.
标准语言模型在维持这种私密的、持续更新的心理状态方面要弱得多。它们最稳定的记忆形式通常是已经生成的token序列。
If the model writes:
如果模型写出:
x=6
x=6
and later writes:
然后写出:
x+3=9,
x+3=9,
those statements remain inside the context. The model can attend to them again when generating the next step.
这些语句保留在上下文中。模型在生成下一步时可以再次关注它们。
Its reasoning is therefore often externalized. The text is not merely a report of a completed thought process. The text is part of the mechanism by which the reasoning occurs.
因此,它的推理通常是外部化的。文本不仅仅是已完成思维过程的报告。文本是推理发生机制的一部分。
Humans do something similar when using scratch paper. The major difference is scale.
人类在使用草稿纸时也会做类似的事情。主要区别在于规模。
An unaided human may struggle to keep five unfamiliar conditions active simultaneously. An AI can preserve dozens or hundreds of them in explicit form.
未经辅助的人类可能难以同时保持五个不熟悉的条件。而人工智能可以以显式形式保留数十或数百个条件。
This does not mean that every item in a long context is retrieved perfectly. Models can overlook relevant information, become distracted or lose track of details. Advertised context length is not the same as perfectly usable memory.
这并不意味着长上下文中的每一项都能被完美检索。模型可能会忽略相关信息、分心或丢失细节。广告宣传的上下文长度并不等同于完全可用的记忆。
Nevertheless, the difference in potential capacity is enormous.
然而,潜在容量的差异是巨大的。
Why this matters particularly for mathematics
为什么这对数学尤其重要
The context-window advantage is not equally useful in every kind of reasoning.
上下文窗口的优势并非在所有类型的推理中都同样有用。
It matters especially for mathematics because mathematical reasoning can be translated unusually well into explicit symbols.
它对数学尤其重要,因为数学推理可以异常良好地转化为显式符号。
Almost every relevant element of a mathematical problem can be written down:
数学问题中几乎所有相关元素都可以被写下来:
- the assumptions;
- 假设;
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力