Anthropic实验室:Claude自主发现新型ART酶系统
Claude discovers a novel enzyme system with CRISPR-like repeats
这是AI Agent从“辅助工具”迈向“独立发现者”的关键里程碑,直接验证了通用大模型在基础科学领域的自主探索能力,值得所有关注AI for Science的团队深入研读其工作流设计。
Science
科学
Claude discovers a novel enzyme system with CRISPR-like repeats
Claude 发现了一种具有类似 CRISPR 重复序列的新型酶系统
Sep 23, 2026
2026年9月23日
We’re introducing a new life sciences research group and laboratory at Anthropic. Our focus is on fundamental biology research using Claude: exploring datasets of DNA to identify uncharacterized protein families, generating hypotheses at scale, and testing them through experiments in the lab. This post introduces the team behind this work and shares early results in which Claude discovered a novel enzyme system with properties reminiscent of CRISPR, with only high-level direction from our scientists.
我们宣布在 Anthropic 成立一个新的生命科学研究小组和实验室。我们的重点是利用 Claude 进行基础生物学研究:探索 DNA 数据集以鉴定未表征的蛋白质家族,大规模生成假设,并通过实验室实验对其进行测试。本文介绍了从事这项工作的团队,并分享了早期成果,其中 Claude 在科学家仅提供高层指导的情况下,发现了一种具有类似 CRISPR 特性的新型酶系统。
Many discoveries that have revolutionized biology and medicine started with a scientist noticing something odd in the staggering diversity of molecular machines found in nature. Restriction enzymes, proteins that cut DNA at specific short sequences, were found in bacterial immune systems, where they destroy the DNA of invading viruses. Researchers realized they could use these enzymes to cut DNA at chosen places and splice genes from one organism into another, which launched the biotechnology industry. Taq polymerase, an enzyme that copies DNA at high temperatures, was identified in a bacterium in a Yellowstone hotspring. It became the basis for PCR, the DNA-copying method used in much of modern diagnostics. CRISPR was first noticed as an unusual repeat sequence in the DNA of certain bacteria, and is now the foundation of gene editing-based medicines.
许多彻底改变生物学和医学的发现,都始于科学家注意到自然界中发现的分子机器惊人多样性中的某些异常现象。限制性内切酶是一种能在特定短序列处切割 DNA 的蛋白质,最初是在细菌免疫系统中发现的,它们用于摧毁入侵病毒的 DNA。研究人员意识到可以利用这些酶在选定位置切割 DNA,并将一个生物的基因拼接至另一个生物体内,这催生了生物技术产业。Taq 聚合酶是一种在高温下复制 DNA 的酶,它是在黄石国家公园温泉中的一种细菌中被识别出来的。它成为了 PCR(聚合酶链式反应)的基础,而 PCR 是现代诊断中广泛使用的 DNA 复制方法。CRISPR 最初是在某些细菌的 DNA 中被发现的一种不寻常的重复序列,如今已成为基于基因编辑药物的基础。
In the spring of 2026, we formed a research group to see whether general AI models can systematize and accelerate such discoveries. We believe that this acceleration will come from establishing a new way of doing biology research, in which agents collaborate with humans in every step of the process. Developing this new way of working required that we build our own lab and a single team working on everything from training Claude in biology to running experiments in the lab.
2026 年春季,我们组建了一个研究小组,旨在考察通用 AI 模型是否能够系统化并加速此类发现。我们相信,这种加速将来自于建立一种新的生物学研究方法,即智能体(agents)与人类在流程的每一步中进行协作。开发这种新的工作方式要求我们建立自己的实验室,并组建一个单一团队,负责从训练 Claude 的生物学生知识到运行实验室实验的所有工作。
Today, we’re sharing early results from one of our first research programs, in which Claude autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR. Although we don’t yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems, all of which are programmable and perform operations like cutting, copying, and pasting DNA. Beyond CRISPR, which has already transformed science and medicine, several other such systems are now in development as promising tools.
今天,我们分享了我们在早期研究项目之一中取得的初步成果。在该项目中,Claude 自主发现了一种新型酶系统,该系统与一系列 DNA 重复序列相关,其模式令人联想到 CRISPR。尽管我们尚不清楚其功能,但 Claude 发现的这套系统具有一组特征,这些特征此前仅在少数其他系统中共同出现过,而这些系统均具备可编程性,并能执行切割、复制和粘贴 DNA 等操作。除了已经彻底改变科学和医学领域的 CRISPR 之外,目前还有几种类似的系统正在开发中,被视为极具前景的工具。
The system that Claude found is based on a reverse transcriptase (RT), enzymes that copy RNA into DNA. While this underlying RT, found in a jumbo phage, had been identified in previous studies, Claude appears to be the first to notice the system’s defining features—an associated array of non-coding DNA sequences and an additional accessory protein of unknown function.
Claude 发现的这套系统基于逆转录酶(RT),这是一种将 RNA 复制为 DNA 的酶。虽然此前研究中已鉴定出这种源自巨型噬菌体的基础 RT,但 Claude 似乎是第一个注意到该系统定义性特征的——即与之关联的非编码 DNA 序列阵列以及一个功能未知的辅助蛋白。
After reviewing the pre-print, Feng Zhang, one of the pioneers of CRISPR genome editing and a professor at MIT and the Broad Institute said:
在审阅了这篇预印本后,CRISPR 基因组编辑的先驱之一、麻省理工学院(MIT)和 Broad 研究所教授 Feng Zhang 表示:
This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation. I hope this work encourages more scientists to explore how AI can support their research.
这是一个令人振奋的例子,展示了 AI 智能体如何助力生物发现。识别与逆转录酶相关的 RNA-重复序列阵列确实引人入胜,值得进一步调查。我希望这项工作能鼓励更多科学家探索如何利用 AI 支持他们的研究。
We gave Claude a prompt to search through a massive database of DNA sequences for interesting new examples of RTs. Our involvement was limited to the initial prompt and the lab work, while Claude agents combed through the database, investigated the distinct RT families, and used their own judgement to identify interesting candidates. After 21 hours spent searching this data by roughly 950 agents using 210 million tokens, one of the agents spotted something remarkable: a repeating pattern of DNA sequences that occurs next to the gene for an odd-looking RT. After further analysis and testing in our lab, we recognized that this pattern marked a previously uncharacterized enzyme system found in bacteriophages (the viruses that infect bacteria) that we call array-associated reverse transcriptases (ART).
我们向 Claude 发出提示,要求其在庞大的 DNA 序列数据库中搜索有趣的新型逆转录酶实例。我们的参与仅限于初始提示和实验室工作,而 Claude 智能体则遍历数据库,调查不同的 RT 家族,并运用自身判断力识别出有趣的候选对象。经过约 950 个智能体耗时 21 小时、消耗 2.1 亿个 token 的数据搜索后,其中一个智能体发现了非凡之处:在一组外观奇特的逆转录酶基因旁边,存在一种重复出现的 DNA 序列模式。经过进一步的分析和实验室测试,我们认识到该模式标记了一种此前未被表征的、存在于噬菌体(感染细菌的病毒)中的酶系统,我们将其称为阵列相关逆转录酶(ART)。
Our work to understand the primary function of ARTs is ongoing. However, we think it is important to share such findings early, both to demonstrate Claude’s capabilities and to give the broader community insight into what we’re working on. We have released a pre-print (here) that discusses this in more detail.
我们了解 ARTs 主要功能的工作仍在进行中。然而,我们认为尽早分享这些发现很重要,这既能展示 Claude 的能力,也能让更广泛的社区深入了解我们正在开展的工作。我们已经发布了一份预印本(此处),对此进行了更详细的讨论。
About our lab
关于我们的实验室
We are a team of scientists who have spent our careers exploring unusual proteins, and specialize in using computational approaches to systematically read DNA, interpret its evolution, and pick out biological systems for further characterization. Our research prior to joining Anthropic has helped to better understand the evolution and regulation of CRISPR systems, discover new enzymes for next-generation cell and gene therapies, and build tools for accelerating the identification of anomalies in DNA, such as human pathogenic variants. We are part of Anthropic’s life sciences organization, alongside teams whose work includes drug discovery, and training Claude in biology and chemistry.
我们是一支科学家团队,职业生涯致力于探索非典型蛋白质,并专长于利用计算方法系统地解读 DNA、阐释其进化过程,并筛选出需要进一步表征的生物系统。在加入 Anthropic 之前,我们的研究有助于更好地理解 CRISPR 系统的进化与调控,发现用于下一代细胞和基因疗法的新酶,并构建工具以加速识别 DNA 中的异常现象,例如人类致病性变异。我们是 Anthropic 生命科学组织的一部分,与该组织中从事药物发现以及训练 Claude 生物学和化学知识的其他团队共同工作。
Our lab, located in the Bay Area, looks like a typical molecular biology lab. We do research that involves only the lower-levels of the biosafety risk level (BSL-1 and BSL-2) and we do not handle pathogens that can infect humans. All of the lab work is performed by human scientists. Although we’ve experimented with using AI to accelerate lab work with initiatives like the Model Hardware Standard, this approach is less conducive to the sort of ad hoc workflows that are involved in our molecular biology research.
我们位于湾区的实验室看起来像是一个典型的分子生物学实验室。我们的研究仅涉及生物安全等级较低的部分(BSL-1 和 BSL-2),我们不处理能感染人类的病原体。所有实验室工作均由人类科学家执行。虽然我们曾尝试通过“模型硬件标准”(Model Hardware Standard)等举措利用 AI 加速实验室工作,但这种方法不太适合我们分子生物学研究中涉及的临时性工作流。
How we work
我们的工作方式
Many of our workflows involve having Claude search through the vast collection of DNA sequences associated with proteins without a known function. One typical pattern begins with a survey of a given protein family. Claude reads the relevant literature and reproduces the established results from public data to check its methods. It then searches for family members or genomic neighbors that fit no described system, and writes a short, human-readable report for each candidate that proposes a function and describes the evidence supporting its claims. In follow-up analyses, Claude critically evaluates the evidence—typically most candidates are eliminated at this stage. A survey may end with a single candidate worth testing, or with none.
我们的许多工作流程都涉及让 Claude 搜索与未知功能蛋白质相关的海量 DNA 序列集合。一种典型的模式始于对特定蛋白质家族的调查。Claude 会阅读相关文献,并利用公共数据复现既定结果以验证其方法。随后,它会寻找不属于任何已描述系统的家族成员或基因组邻近区域,并为每个候选者撰写一份简短、易于人类阅读的报告,提出其功能假设并描述支持该主张的证据。在后续分析中,Claude 会对证据进行批判性评估——通常大多数候选者在此阶段被排除。一次调查可能最终得出一个值得测试的候选者,也可能一无所获。
When a candidate survives our review, we test it in the laboratory, expressing the protein in standard laboratory strains and characterizing it biochemically and structurally, with Claude helping to interpret the data. We do our work in Claude Science and Claude Code, the same tools available to any scientist, and sometimes with a harness of our own that coordinates many Claude sessions running in parallel.
当候选方案通过我们的审查后,我们会在实验室中进行测试,在标准实验室菌株中表达该蛋白,并从生化与结构层面对其进行表征,同时借助 Claude 来解读数据。我们在 Claude Science 和 Claude Code 中开展工作——这些是任何科学家均可使用的工具,有时还会配合我们自行开发的协调框架,以并行运行多个 Claude 会话。
Because Claude produces hypotheses so prolifically, the hypotheses themselves have become an object of study for us. With hundreds to thousands of candidate reports from a single campaign, we have been asking what distinguishes the proposals we judge worth testing from those we set aside. What we learn goes back into the instructions we give Claude and teaches it to mimic our own scientific taste.
由于 Claude 能大量生成假设,这些假设本身已成为我们研究的对象。在一次活动中会产生数百至数千份候选报告,我们一直在探究:那些我们认为值得测试的提案,与那些被搁置的提案之间有何区别?我们从中学到的经验会反馈到给 Claude 的指令中,并教会它模仿我们自身的科学品味。
Claude finds ART
Claude 发现 ART
In the past few years, researchers have discovered many more reverse transcriptases (RTs), most of them in bacteria, where they act as part of the immune system. Nearly all RT families were found by genomic analysis, or genome mining, which requires researchers to search sequence databases for genes that no one has characterized, notice the unusual ones, and work out what they do.
在过去几年里,研究人员发现了更多逆转录酶(RT),其中大多数存在于细菌中,作为免疫系统的一部分发挥作用。几乎所有 RT 家族都是通过基因组分析(或称基因组挖掘)发现的,这需要研究人员在序列数据库中搜索尚未被表征的基因,识别出其中的异常者,并阐明其功能。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力