跳到主内容
@wquguru
精选88NVIDIA 博客(RSS)模型发布/更新

NVIDIA联合DeepMind开源2800种病毒蛋白复合物结构数据

How Open Science Can Help Researchers Prepare for the Next Pandemic

原文
发到 X
推荐理由

AI for Science领域的重磅开源,直接提供2800+病毒蛋白复合物的高置信度结构数据,对生物医药研发极具实用价值,建议相关领域研究者重点关注。

When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time. The next pandemic may not offer the same head start.

当 COVID-19 出现时,科学家们拥有一个关键优势:数十年来对冠状病毒的先前研究使他们充分了解该病毒的关键蛋白质,从而能够在创纪录的时间内设计出疫苗。下一次大流行可能不会提供同样的先发优势。

To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google DeepMind and the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI), to release predicted 3D structures for the protein complexes of more than 2,800 viruses — openly available to any scientist, anywhere, through the AlphaFold Database.

为了提高胜算,NVIDIA 加入了一个由全球研究机构组成的联盟,包括 Google DeepMind 和欧洲分子生物学实验室的欧洲生物信息学研究所(EMBL-EBI),通过 AlphaFold Database 向世界各地的任何科学家开放发布超过 2,800 种病毒的蛋白质复合物预测三维结构。

The structures in the newly released dataset were inferred using AlphaFold2 — Google DeepMind’s AI model for predicting how proteins fold into 3D shapes — with optimization from NVIDIA BioNeMo Inference Runtime. This allowed the team to scale inference to thousands of viral proteomes, predicting the complexes, or groups of interacting proteins, encoded within each virus.

新发布数据集中的结构是使用 AlphaFold2 推断得出的——这是 Google DeepMind 用于预测蛋白质如何折叠成三维形状的 AI 模型——并经过 NVIDIA BioNeMo Inference Runtime 优化。这使得团队能够将推理扩展到数千个病毒蛋白质组,预测每个病毒中编码的复合物(即相互作用的蛋白质组)。

“Our ambition with the AlphaFold Database has always been to democratize access to foundational biology at scale,” said Risha Patel, life sciences partnerships manager at Google DeepMind. “This collaboration to bring thousands of viral complexes into the database will equip scientists around the world with insights they need to help prepare for future outbreaks.”

Google DeepMind 生命科学合作伙伴关系经理 Risha Patel 表示:“我们构建 AlphaFold Database 的愿景始终是以规模化方式普及基础生物学知识的获取。”“此次合作将数千种病毒复合物引入数据库,将为世界各地的科学家提供他们所需的见解,以帮助为未来的疫情爆发做好准备。”

NVIDIA is also openly releasing the BioNeMo Structure Prediction Pipeline, the GPU-accelerated workflow used to generate the dataset, so researchers can go from protein sequence to predicted 3D structure for their own targets.

NVIDIA 还公开发布了 BioNeMo Structure Prediction Pipeline(BioNeMo 结构预测管道),这是用于生成数据集的 GPU 加速工作流程,研究人员可以利用它从蛋白质序列生成其自身目标的预测三维结构。

Preparation for the next pandemic must begin now. An analysis by the Center for Global Development estimates a roughly 50% chance of the world facing a pandemic as severe as COVID-19 by 2050.

必须从现在开始为下一次大流行做准备。全球发展中心(Center for Global Development)的一项分析估计,到 2050 年,世界面临与 COVID-19 严重程度相当的大流行的概率约为 50%。

“When the next pandemic happens, there may be something that comes out of the blue, and we’ll be lacking the knowledge we had for COVID,” said Joe Grove, professor of molecular virology at the Medical Research Council-University of Glasgow Centre for Virus Research and a collaborator on the project. “What we’re trying to do is stockpile some of that knowledge ahead of time.”

格拉斯哥大学医学研究委员会病毒研究中心的分子病毒学教授、该项目合作者 Joe Grove 表示:“当下一次大流行发生时,可能会出现一些突如其来的情况,而我们可能会缺乏在 COVID 时期所拥有的知识。”“我们要做的是提前储备部分这些知识。”

About 30% of the protein interactions being added to the database are completely new to science, showing interaction shapes that have never been documented in the Protein Data Bank, the main repository of experimentally determined protein structures. This translates to new insights for the biological community to explore and harness to generate new knowledge.

添加到数据库中的蛋白质相互作用中,约30%对科学界而言是完全全新的,展示了此前在蛋白质数据库(Protein Data Bank,即实验确定的蛋白质结构的主要存储库)中从未记录过的相互作用形态。这为生物学界提供了新的见解,可供其探索并利用以生成新知识。

“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA. “We’re enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes, so the whole field can move forward.”

"该数据库是假设生成的引擎,"英伟达数字生物学应用研究科学团队负责人Chris Dallago表示。"我们使生物学家和人工智能社区能够调查蛋白质相互作用,不仅作为单个分子,而且作为复合物,从而使整个领域得以前进。"

Predicting Complex Protein Structures

预测复杂蛋白质结构

Most proteins don’t work alone — they come together in complexes of multiple molecules to perform sophisticated functions. Those structures are often what a vaccine or drug must target to disrupt viral function.

大多数蛋白质并非单独工作——它们以多个分子的复合物形式聚集在一起,以执行复杂的功能。这些结构通常是疫苗或药物必须靶向以破坏病毒功能的目标。

Understanding the 3D structure of the COVID-19 virus’ spike protein, for example, proved foundational to vaccine design. For thousands of other viruses, no such structural knowledge exists today. This dataset begins to fill that gap.

例如,了解新冠病毒刺突蛋白的三维结构被证明是疫苗设计的基础。对于其他数千种病毒,目前尚不存在此类结构知识。该数据集开始填补这一空白。

Traditional methods for determining protein structures — crystallizing proteins and shooting X-rays at them — can take years and cost thousands of dollars per structure. AlphaFold2, which was optimized with NVIDIA BioNeMo to efficiently run on NVIDIA GPUs, predicts a structure in minutes and can be run in bulk. Scientists can then verify high-confidence predictions through experimental methods.

确定蛋白质结构的传统方法——结晶蛋白质并向其发射X射线——可能需要数年时间和每结构数千美元的成本。经过英伟达BioNeMo优化以高效运行在英伟达GPU上的AlphaFold2可以在几分钟内预测出结构,并且可以批量运行。然后,科学家可以通过实验方法验证高置信度的预测结果。

For this project, the team systematically worked through the protein structures of viral families known to infect humans, from common-cold viruses to emerging threats like Mpox.

对于这个项目,团队系统地研究了已知会感染人类的病毒家族的蛋白质结构,从普通感冒病毒到像猴痘这样的新兴威胁。

A Global Collaboration With Global Access

全球协作与全球访问

The collaboration spans the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.

此次合作涵盖了流行病防范创新联盟、欧洲分子生物学实验室欧洲生物信息学研究所(EMBL-EBI)、Google DeepMind、英伟达、首尔国立大学、成均馆大学、瑞士生物信息学研究所和格拉斯哥大学。

The dataset release — coinciding with a United Nations General Assembly meeting convened by the World Economic Forum on pandemic prevention, preparedness and response taking place this week in New York City — contributes to the AlphaFold Database, which now holds more than 260 million protein and protein complex predictions covering nearly every cataloged protein known to science.

该数据集的发布恰逢世界经济论坛本周在纽约市召集联合国大会会议,讨论大流行的预防、准备和应对,并为AlphaFold数据库做出贡献,该数据库现在拥有超过2.6亿个蛋白质和蛋白质复合物的预测,涵盖了科学界已知的几乎每一种目录化蛋白质。

“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL-EBI. “The dataset also covers lesser-studied viruses and lowers the barriers for scientists in low-resource settings who are confronting outbreaks firsthand.”

“公开这些数据对于理解病毒诊断以及开发治疗方法与疫苗至关重要,”EMBL-EBI(欧洲生物信息学研究所)临时主任Jo McEntyre表示。“该数据集还涵盖研究较少的病毒,并降低了资源匮乏地区一线应对疫情的科学家的使用门槛。”

Predictions in the open dataset are labeled by confidence. The structures show what viral complexes may look like and how individual proteins might interact within a viral proteome.

开放数据集中的预测结果标注了置信度。这些结构展示了病毒复合物可能的外观形态,以及单个蛋白质在病毒蛋白质组中可能的相互作用方式。

Overall, the new data represents a major contribution to the information available for scientists across digital biology and disease research.

总体而言,新数据为数字生物学和疾病研究领域的科学家提供了重要的信息贡献。

“When I did my Ph.D., there were no structures for any of the proteins we were investigating. It was like working in the dark — we had to guess what was going on,” said Grove. “This dataset is a powerful tool for all the researchers doing their Ph.D.s now, giving them high-quality structural data that’s going to accelerate fundamental science.”

“当我攻读博士学位时,我们研究的任何蛋白质都没有结构数据。那就像在黑暗中摸索——我们必须猜测正在发生的事情,”Grove说。“该数据集是如今所有正在进行博士研究的科研人员的强大工具,为他们提供高质量的结构性数据,这将加速基础科学的进展。”

Explore the viral protein complex dataset on the AlphaFold Database Pandemic Preparedness Portal, predict structures for protein targets with the BioNeMo Structure Prediction Pipeline, and learn more about NVIDIA BioNeMo.

在AlphaFold数据库大流行准备门户上探索病毒蛋白质复合物数据集,使用BioNeMo结构预测管道预测蛋白质靶点的结构,并了解更多关于NVIDIA BioNeMo的信息。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件