AI 推动科学进步需推理而非仅靠数据
AI for science needs reasoning, not just data
做 AI for Science 的同学必看,这篇把 AlphaFold 式数据瓶颈和智能体路线讲透了,还给了 Co-Scientist 的完整工作流,值得收藏对照自己的科研场景。
Every few decades, someone announces that science has reached its end. In 1903, the revered physicist Albert Michelson wrote that the “facts of physical science have all been discovered.” In the 1980s, Stephen Hawking predicted that theoretical physics might be finished by the end of the century. With the explosive arrival of artificial intelligence, the feeling is in the air again—this time accompanied by a Nobel Prize.
每隔几十年,就会有人宣布科学已经走到了尽头。1903年,备受尊敬的物理学家阿尔伯特·迈克尔逊写道,'物理科学的事实都已被发现。'在1980年代,斯蒂芬·霍金预测理论物理学可能在本世纪末结束。随着人工智能的爆炸性到来,这种感觉再次弥漫在空气中——这一次还伴随着诺贝尔奖。
In 2024, Demis Hassabis and John Jumper of Google DeepMind were awarded part of the Nobel in chemistry for their neural network AlphaFold, which predicts the three-dimensional structures of proteins by learning from thousands of experimentally measured shapes. This devilish problem had resisted systematic attacks for half a century; AlphaFold seemed to have solved it once and for all, and the world became fixated on the promise of its approach. Hassabis and his team called AlphaFold “the template for how AI can accelerate all of science to digital speed.” A wave of startups building foundation models for biology, chemistry, and materials discovery raised billions of dollars, buoyed by DeepMind’s success. AlphaFold had shown that the combination of AI and sufficient data could make groundbreaking discoveries (even if we did not understand the underlying mechanisms involved), and it seemed, once again, that a path through the rest of science was laid out before us.
2024年,谷歌DeepMind的德米斯·哈萨比斯和约翰·江珀因他们的神经网络AlphaFold获得了诺贝尔化学奖的一部分,该网络通过从数千个实验测量的形状中学习来预测蛋白质的三维结构。这个棘手的问题半个世纪以来一直抵抗着系统性的攻击;AlphaFold似乎一劳永逸地解决了它,世界也因此对其方法的潜力着迷。哈萨比斯和他的团队称AlphaFold为'AI如何将整个科学加速到数字速度的模板'。一批为生物学、化学和材料发现构建基础模型的初创公司,在DeepMind成功的推动下筹集了数十亿美元。AlphaFold已经表明,AI与足够数据的结合可以产生突破性的发现(即使我们不了解所涉及的底层机制),而且似乎又一次,一条通往科学其余部分的道路展现在我们面前。
To be sure, AI will bring extraordinary changes to science, but it has become increasingly clear that AlphaFold, and things like it, may not be the best template for that metamorphosis. Though it is a profound achievement, the conditions that produced the likes of AlphaFold are rare, and the time it will take to meet those conditions in other fields will be measured in decades, not years. Instead, the acceleration of science will come about thanks to another approach: AI agents.
诚然,AI将给科学带来非凡的变化,但越来越清楚的是,AlphaFold及其类似物可能不是这种变革的最佳模板。尽管它是一个深远的成就,但产生像AlphaFold这样的条件很少见,而在其他领域满足这些条件所需的时间将以数十年计,而不是数年。相反,科学的加速将归功于另一种方法:AI代理。
The primary condition for AlphaFold’s success was the existence of the Protein Data Bank, a data set of roughly 170,000 experimentally validated protein structures on which DeepMind’s team could train its model. The creation of the Protein Data Bank was not simple: It took 53 years of international scientific cooperation and, by a recent estimate, roughly $21 billion worth of experimental work to assemble. Efforts of that scale are infamously difficult to fund, next to impossible to coordinate, and hugely time-consuming to execute; they have often been unsuccessful as a result.
AlphaFold成功的主要条件是蛋白质数据银行的存在,这是一个包含约17万个实验验证的蛋白质结构的数据集,DeepMind团队可以在此基础上训练他们的模型。蛋白质数据银行的创建并不简单:它花费了53年的国际科学合作,并且根据最近的估计,大约需要210亿美元的实验工作来组装。这种规模的努力众所周知难以资助,几乎不可能协调,执行起来也非常耗时;因此它们常常不成功。
But even in fields with the requisite cohesion and resources, and where the relevant data are not rendered inaccessible by commercial ownership, another barrier is too little discussed: the scientific impossibility of generating comparable data. In the case of protein structures, the key experimental technique—protein crystallography—is an unusually replicable and dependable tool, so much so that over 25 Nobel Prizes have relied on it. But in most of experimental science, results vary more often than not. Cell lines drift. Chemicals have trace contaminants. Lab humidity changes. The creation of measured datasets that will be consistent enough, accurate enough, precise enough, and scalable enough to train a modern neural network in biology or most of chemistry would require new kinds of measurement and new standardized approaches—none of which will be ready anytime soon.
但即使在具备必要凝聚力和资源的领域,且相关数据未被商业所有权封锁的情况下,另一个障碍也很少被讨论:生成可比数据在科学上的不可能性。就蛋白质结构而言,关键实验技术——蛋白质晶体学——是一种异常可重复且可靠的工具,以至于超过25项诺贝尔奖依赖于它。但在大多数实验科学中,结果往往不尽相同。细胞系会漂移。化学品含有微量污染物。实验室湿度会变化。要创建足以训练现代神经网络(在生物学或大部分化学领域)的测量数据集,需要足够一致、准确、精确且可扩展的数据,这将需要新型测量方法和新的标准化方法——而这些在短期内都无法准备就绪。
Of course, there are a handful of fields where these requirements are met: weather forecasting, much of genomics, very limited areas of chemistry. These may see AlphaFold-style breakthroughs soon, if they haven’t already. Government support for the production and coordination of those datasets will be critical, as the US National Security Commission on Emerging Biotechnology has argued. But for most open questions in science, we will need a different plan, at least in the short term. Luckily, something quieter and more modest has begun to show promise.
当然,有一些领域满足这些要求:天气预报、大部分基因组学、非常有限的化学领域。这些领域可能很快就会迎来AlphaFold式的突破,如果还没有的话。政府对这些数据集的生产和协调的支持将至关重要,正如美国国家新兴生物技术安全委员会所主张的那样。但对于科学中的大多数未解问题,我们至少短期内需要不同的计划。幸运的是,一种更安静、更谦逊的方法已开始显示出前景。
Scientists have always reasoned under uncertainty. Biologists working to identify new drug targets have never had perfect datasets. Instead, they combine docking calculations and known structures, factor in molecular dynamics, run a handful of binding assays, and use their judgment to weigh each method according to its particular strengths and points of failure. The skill of science is not in any single tool; it is synthesizing what many tools produce, and revising the results as the evidence comes in. This is how most working research actually proceeds. But until very recently, no software could do it.
科学家们一直在不确定性下推理。致力于识别新药物靶点的生物学家从未拥有完美的数据集。相反,他们结合对接计算和已知结构,考虑分子动力学,进行少量结合实验,并根据每种方法的特定优势和失败点,运用判断力权衡每种方法。科学的技能不在于任何单一工具;而在于综合多种工具产生的成果,并根据证据不断修正结果。这就是大多数实际研究工作的进行方式。但直到最近,还没有软件能够做到这一点。
Agents now can. Simply put, an agent is an AI reasoning engine that has been given access to tools—digital or physical—and the capabilities to use them. Over the last few years, a fundamental architectural shift in AI has enabled the rapid proliferation of these programs, which are powered by large language models, dramatically reducing the need for scientifically specialized datasets. For science, this technological advancement represents a foundational change: it has allowed us to create digital tools that can mimic the iterative, highly contingent process of actual research. While tools like AlphaFold apply a powerful approach to a limited question, agents are inherently generalists. They do not represent a new way to do science—instead, they digitally model the human process of discovery.
智能体现在可以做到。简而言之,智能体是一个被赋予工具(数字或物理)及使用这些工具的能力的AI推理引擎。在过去几年中,AI架构的根本性转变使得这些由大型语言模型驱动的程序迅速普及,大大减少了对科学专用数据集的需求。对科学而言,这一技术进步代表了一种根本性的变革:它使我们能够创建模拟实际研究过程中迭代且高度偶然性的数字工具。虽然像AlphaFold这样的工具将强大的方法应用于有限的问题,但智能体本质上是通才。它们并不代表一种新的科学研究方式——而是以数字方式模拟人类的发现过程。
Consider Google’s AI Co-Scientist, announced in May. Researchers gave it a one-page brief and a goal: Figure out how antibiotic resistance spreads between bacterial species, a key driver of drug-resistant infections. The system spun up sub-agents. One drafted hypotheses from the literature. Another picked them apart like a peer reviewer. A third ran tournaments to rank the strongest candidates. A fourth refined the winning hypothesis. The agent concluded that resistance genes were hitching rides on bacterial viruses, borrowing whichever virus could ferry them into a new host. The hypothesis was correct. Researchers at Imperial College London had spent a decade reaching the same conclusion through painstaking wet-lab work; their paper, previously unseen by Co-Scientist, was still in peer review.
以谷歌于五月发布的AI联合科学家为例。研究人员给了它一页简报和一个目标:找出抗生素耐药性如何在细菌物种之间传播,这是耐药性感染的关键驱动因素。该系统生成了子智能体。一个从文献中起草假设。另一个像同行评审员一样对其进行剖析。第三个运行锦标赛以对最强候选进行排名。第四个完善获胜的假设。该智能体得出结论:耐药基因搭上了细菌病毒的便车,借用任何能将它们运送到新宿主的病毒。这个假设是正确的。伦敦帝国理工学院的研究人员花了十年时间通过艰苦的湿实验得出了同样的结论;他们的论文此前未被联合科学家看到,仍在同行评审中。
Agents like Co-Scientist are still novel tools, and there are real challenges to overcome before they become a ubiquitous part of the scientific process: They are still liable to hallucinate, their judgment is not consistent, and they have memory and input constraints that limit the time they can run autonomously. But these technical barriers will fall away, and as they do we will begin to notice the compounding effects of scientific agents on the reliability, consistency, and velocity with which science is done.
像联合科学家这样的智能体仍然是新颖的工具,在它们成为科学过程中普遍存在的一部分之前,还有真正的挑战需要克服:它们仍然容易产生幻觉,判断力不一致,并且存在记忆和输入限制,限制了它们自主运行的时间。但这些技术障碍将会消失,随着它们的消失,我们将开始注意到科学智能体对科学研究的可靠性、一致性和速度所产生的复合效应。
Perhaps most notably, agents offer a structural fix for science’s “reproducibility crisis,” the widespread problem of researchers’ inability to replicate each other’s results. For decades, the scientific community has begged researchers to share their raw data and exact code in an effort to standardize experimental processes. But researchers have long resisted this tedious administrative work, which happens after the interesting science is already done. Agents, in contrast, automatically log every move they make, creating an exact record of the method that led to their results and allowing for precise replication.
也许最值得注意的是,智能体为科学的“可重复性危机”提供了一种结构性修复方案,即研究人员无法复制彼此结果的普遍问题。几十年来,科学界一直恳求研究人员分享他们的原始数据和确切代码,以努力标准化实验过程。但研究人员长期以来一直抵制这种繁琐的行政工作,这些工作发生在有趣的科学完成之后。相比之下,智能体会自动记录它们的每一个动作,创建导致其结果的方法的精确记录,并允许精确复制。
A second consequence will be an amplification of scientific memory. The transfer of knowledge between researchers is a famously murky process; if it isn’t done over years of training and observation, graduate students are left to pore through the messy lab notebooks kept by decades of predecessors, looking for the details that will make or break their protocol. As agents become an increasingly large part of the scientific process, though, a lab’s entire scientific history will be recorded in a central, standardized repository of institutional knowledge.
第二个后果将是科学记忆的增强。研究人员之间的知识转移是一个出了名的模糊过程;如果不是通过多年的培训和观察来完成,研究生们只能翻阅几十年前前辈留下的杂乱实验室笔记本,寻找那些决定他们实验方案成败的细节。然而,随着智能体在科学过程中占据越来越大的比例,实验室的整个科学历史将被记录在一个集中的、标准化的机构知识库中。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力