跳到主内容
@wquguru
精选75MIT Technology Review AI(RSS)模型发布/更新

AI 模型在这些智力测试中翻车,你能答对吗?

AI models flub these intelligence tests. Can you fare any better?

原文
发到 X

Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the Chinese board game Go are famous AI test beds too.

谜题和游戏从一开始就是人工智能发展的核心。正如我们人类喜欢用填字游戏或逻辑谜题来测试自己的智慧一样,开发者也可以通过一系列游戏挑战来测试模型的进步程度。“机器学习”一词在1959年由IBM计算机科学家Arthur Samuel发表的一篇关于学习下棋算法的文章中得以普及。国际象棋和中国围棋也是著名的AI测试平台。

Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time.

仅从解谜能力来看,AI正在大幅且快速地提升。2024年底,哥伦比亚大学的科学家团队表明,即使是最优秀的模型也只能解开臭名昭著的《纽约时报》Connections谜题的18%;到2025年初,一些模型几乎每次都能完美解开这些谜题。

But puzzles do more than just highlight the inexorable advance of AI capabilities. Seeing where models succeed and fail—and where we humans still beat them—can provide a useful window into the technology’s strengths and weaknesses. Despite advances, today’s models still fumble: Subtle changes in classic riddles often trip them up, and visual puzzles are a particular weak spot.

但谜题不仅仅凸显了AI能力不可阻挡的进步。观察模型在哪些地方成功、哪些地方失败——以及我们人类在哪些地方仍能胜过它们——可以为技术的优缺点提供一个有用的窗口。尽管取得了进步,今天的模型仍然会犯错:经典谜题中的细微变化常常会让它们出错,而视觉谜题则是一个特别薄弱的环节。

Here you’ll have the chance to test your wits on puzzles that have stumped models at one time or another. Some might be as tricky for you as they were for the AI; others are so simple that they’ll have you doubting whether AI is really intelligent at all. Each one highlights at least one way in which machine and human cognition differ. If you ace the test, you’ll have proved that you can out-puzzle an AI—at least for now.

在这里,你将有机会用那些曾难倒模型的谜题来测试自己的智慧。有些谜题对你来说可能和AI一样棘手;另一些则简单到让你怀疑AI是否真的智能。每个谜题都至少突显了机器和人类认知的一种差异。如果你通过了测试,你就证明了自己能在解谜上胜过AI——至少目前是这样。

Spatial Reasoning

空间推理

Let’s start with a domain where humans have a huge advantage: spatial reasoning. If you’ve ever taken an IQ test, you may have done a mental rotation problem. These puzzles ask you to determine whether different images represent the same objects from different angles. Though today’s language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles. For all the talk of how world models can help AI understand physical environments, LLMs still don’t seem to be able to manipulate 3D objects the way spatial thinkers like architects and mechanical engineers can.

让我们从一个人类拥有巨大优势的领域开始:空间推理。如果你参加过智商测试,你可能做过心理旋转问题。这些谜题要求你判断不同图像是否代表从不同角度看到的同一物体。尽管今天的语言模型通常具备分析视觉输入的能力,但它们在解决这些谜题时仍然表现糟糕。尽管有关于世界模型如何帮助AI理解物理环境的种种讨论,LLM似乎仍然无法像建筑师和机械工程师这样的空间思考者那样操作3D物体。

Mental Rotation

心理旋转

Instructions: Choose the answer that shows the object in the prompt, but from a different angle. In each case, there’s only one correct answer!

说明:选择显示提示中物体但从不同角度观察的答案。每种情况下只有一个正确答案!

Memory & Adaptability

记忆与适应性

Frontier LLMs have extraordinary memories; they were exposed to a monstrous volume of facts during training and can recite many of them faithfully. That’s an asset for outcompeting humans at trivia, but it can also be a liability. When a puzzle closely resembles one a model saw during training, the model may whiz by key differences and respond with what it memorized.

前沿大型语言模型拥有非凡的记忆力;它们在训练期间接触了海量的事实,并能忠实地复述其中许多内容。这在琐事问答中胜过人类是一项资产,但也可能成为负担。当一道谜题与模型在训练中见过的题目极为相似时,模型可能会忽略关键差异,直接给出它记忆中的答案。

This held true in a 2024 study in which researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of a classic type of puzzle called Knights and Knaves. In these problems, some characters always tell the truth and others always lie, and you have to figure out who’s who. The same principle may be at work in a test called SimpleBench. These questions resemble more complicated problems that models likely encountered in training. Humans spot the trick, but even top-tier models trip.

这一点在2024年的一项研究中得到了证实,该研究中,来自谷歌和伊利诺伊大学厄巴纳-香槟分校的研究人员对一种名为“骑士与无赖”的经典谜题的轻微变体进行了训练和测试。在这些问题中,有些角色总是说真话,有些总是说谎,你需要弄清楚谁是谁。同样的原理可能也适用于一个名为SimpleBench的测试。这些问题与模型在训练中可能遇到的更复杂问题相似。人类能识破其中的陷阱,但即使是顶级模型也会出错。

Knights and Knaves

骑士与无赖

Instructions: The only thing you need to know to solve these puzzles is that knights always tell the truth and knaves always lie. Determine who’s what on the basis of what each character says.

说明:解决这些谜题你只需要知道,骑士总是说真话,无赖总是说谎。根据每个角色所说的话来判断他们是什么身份。

SimpleBench

SimpleBench

Instructions: Read these SimpleBench problems carefully, and you should be able to figure out the answers in no time.

说明:仔细阅读这些SimpleBench问题,你应该很快就能找到答案。

Abstract & Visual Reasoning

抽象与视觉推理

AI doesn’t just bungle visual problems in 3D—two dimensions can trip it up as well. That’s a major factor in how well models do on the most famous ­puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell.

人工智能不仅在3D视觉问题上表现不佳——二维问题同样会让它出错。这是影响模型在最著名的基于谜题的基准测试ARC-AGI上表现的一个主要因素。这些问题要求你从一组示例中推断出抽象的、通用的规则。当模型以数字字符串而非图像的形式接收每个网格(该字符串编码了每个单元格的颜色)时,它们在ARC谜题上的表现会更好。

Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-­generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles—such as the one printed here—still stump them.

研究表明,即使模型正确回答了ARC-AGI问题,它们也常常使用复杂且无法泛化的规则,而人类则依赖简单的视觉概念。尽管存在这些劣势,过去一年中模型在ARC-AGI上的表现已经相当出色,但有些谜题——比如这里印出的这道——仍然会让它们困惑。

ARC-AGI

ARC-AGI

Instructions: Study the three pairs of grids shown below to figure out the rule that dictates how the ones on the left transform into the ones on the right. Then get out your markers or colored pencils and fill in the fourth grid using that rule. (The solution is the same no matter which way the grids are oriented.)

说明:研究下面显示的三对网格,找出决定左侧网格如何转换为右侧网格的规则。然后拿出你的马克笔或彩色铅笔,用该规则填充第四个网格。(无论网格如何旋转,答案都是相同的。)

Intuition

直觉

It’s not just AI models that fall into traps. We humans have our own cognitive foibles, many of which AI does not share. Psychologists have designed problem suites that invert the SimpleBench phenomenon: For these questions, humans often give knee-jerk answers, whereas models will respond deliberatively. Some of the problems exploit errors in the ways that we intuitively do math; others are phrased so as to suggest obvious answers that fall apart if the question is read carefully.

不仅仅是AI模型会落入陷阱。我们人类也有自己的认知缺陷,其中许多是AI所不具备的。心理学家设计了一系列问题套件,这些套件颠倒了SimpleBench现象:对于这些问题,人类常常给出条件反射式的答案,而模型则会深思熟虑地回应。其中一些问题利用了我们在直觉数学运算中的错误;另一些则措辞巧妙,暗示出显而易见的答案,但如果仔细阅读问题,这些答案就会站不住脚。

Lightning Round

闪电回合

Instructions: Answer the questions below as quickly as you can.

说明:尽可能快速地回答以下问题。

Increasing Complexity

复杂度递增

In some cases, whether an LLM can complete a puzzle is a matter of scale. One study from researchers at Apple found that LLMs can ace simple versions of the Tower of Hanoi problem, which involves moving a stack of disks one at a time without ever putting a larger disk atop a smaller one, and river-crossing puzzles, in which a group of people must traverse a river according to certain rules. But only up to a point: As the number of disks or people hits six and higher, the models began to falter.

在某些情况下,LLM能否完成谜题取决于规模。苹果公司研究人员的一项研究发现,LLM能够轻松解决简单版本的汉诺塔问题(涉及一次移动一堆圆盘,且不能将大圆盘放在小圆盘之上)和过河谜题(一群人必须按照特定规则过河)。但仅限于一定程度:当圆盘或人数达到六个及以上时,模型开始出现失误。

In another study, researchers at the University of Washington, Stanford University, and the Allen Institute for AI observed that LLMs struggle similarly with logic grid puzzles, which require deducing the attributes of a set of individuals from a list of clues. The Apple paper went viral, but commentators questioned whether the results reveal a unique limitation of LLM reasoning—or just that it’s normal to make errors as complexity piles up.

在另一项研究中,华盛顿大学、斯坦福大学和艾伦人工智能研究所的研究人员观察到,LLM在逻辑网格谜题上同样表现挣扎,这类谜题要求从一系列线索中推断出一组个体的属性。苹果的论文在网上疯传,但评论者质疑这些结果是否揭示了LLM推理的独特局限性,还是仅仅表明随着复杂度增加,出错是正常现象。

The River

过河

Instructions: Using the scenario provided, plan the trips necessary to get everyone across the river.

说明:根据提供的场景,规划让所有人过河所需的行程。

Logic Grid

逻辑网格

Instructions: Using the list of clues, determine who lives in each house and what style of music each person enjoys. There is only one possible solution. You may find it helpful to fill out the grid below to keep track of your deductions.

说明:根据线索列表,确定谁住在哪栋房子里,以及每个人喜欢哪种音乐风格。只有一个可能的解决方案。填写下面的网格来记录你的推理可能会有所帮助。

Grace Huckins is an AI reporter at MIT Technology Review. They have a PhD in neuroscience.

格蕾丝·哈金斯是《麻省理工科技评论》的AI记者。她拥有神经科学博士学位。

Credits:

致谢:

Mental Rotation: CC BY 4.0. Stogiannidis, Ilias, Steven McDonagh, Sotirios A. Tsaftaris. Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models (copyright 2025); illustrations by John MacNeill. Knights & knaves: Courtesy Dan MacKinnon. Simplebench: CC BY 4.0. SimpleBench Team. The Text Benchmark in which Unspecialized Human Performance Exceeds that of Current Frontier Models (copyright 2024). ARC-AGI: Courtesy ARC Prize Foundation. Lightning round: CC BY 4.0. Hagendorff, Thilo, Sarah Fabi, Michal Kosinski. Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT. Nat Comput Sci 3, 833–838 (copyright 2023). The river: Adapted from Propositiones ad Acuendos Juvenes, Alcuin of York (ca. 800 CE). Logic grid: Apache License 2.0. Lin, Bill Y., Ronan Le Bras, Kyle Richardson, et al. ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning (copyright 2025)

心理旋转:CC BY 4.0。Stogiannidis, Ilias, Steven McDonagh, Sotirios A. Tsaftaris。《注意差距:视觉-语言模型中的空间推理基准测试》(版权2025);插图由John MacNeill绘制。骑士与无赖:由Dan MacKinnon提供。Simplebench:CC BY 4.0。SimpleBench团队。《非专业人类表现超过当前前沿模型的文本基准》(版权2024)。ARC-AGI:由ARC Prize基金会提供。闪电回合:CC BY 4.0。Hagendorff, Thilo, Sarah Fabi, Michal Kosinski。《大型语言模型中出现了类似人类的直觉行为和推理偏差,但在ChatGPT中消失了》。Nat Comput Sci 3, 833–838(版权2023)。河流:改编自《Propositiones ad Acuendos Juvenes》,约克的阿尔昆(约公元800年)。逻辑网格:Apache License 2.0。Lin, Bill Y., Ronan Le Bras, Kyle Richardson, 等。《ZebraLogic:关于LLM逻辑推理的扩展极限》(版权2025)

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近