MIT发布xvr:AI实现术中X光与3D影像秒级配准
New AI technique could make minimally invasive surgeries safer and more precise
医学影像AI落地关键突破,xvr通过物理仿真解决泛化难题,为手术导航提供可复现的高精度方案,值得医疗AI从业者关注。
Researchers created a new technique that accurately and rapidly matches X-rays captured during surgery with a patient’s preoperative 3D medical scan. This method could make it easier for clinicians to precisely pilot minimally invasive surgical tools, leading to faster and safer procedures.
研究人员开发了一种新技术,能够准确且快速地将手术中拍摄的X射线与患者术前的3D医学影像进行匹配。该方法可使临床医生更轻松地精确引导微创手术工具,从而带来更快、更安全的手术过程。
Clinicians perform many minimally invasive surgeries using real-time X-rays to help them steer devices like catheters and endoscopes through tiny incisions. But since X-rays are flat images, it can be challenging to determine exactly where surgical tools are located and oriented within the patient’s body, increasing the risk of complications.
临床医生在进行许多微创手术时,会使用实时X射线来帮助他们通过微小切口引导导管和内窥镜等设备。但由于X射线是平面图像,要确定手术工具在患者体内的确切位置和方向颇具挑战,这增加了并发症的风险。
To help localize surgical devices, clinicians may manually align X-rays with preoperative 3D medical images, such as CT scans or MRIs. Artificial intelligence tools designed to streamline this process struggle to align images robustly for all patients, making them infeasible in practice.
为了帮助定位手术器械,临床医生可能会手动将X射线与术前3D医学影像(如CT扫描或MRI)对齐。旨在简化这一流程的人工智能工具难以在所有患者身上稳健地对齐图像,因此在实践中不可行。
This new system, developed by scientists and clinicians at MIT and collaborating institutions, uses an AI model that adapts to each patient in only about five minutes. The model automatically matches one patient’s X-rays with 3D scans in a matter of seconds, and with sub-millimeter precision.
这个由麻省理工学院(MIT)及合作机构的科学家和临床医生共同开发的新系统,使用了一个AI模型,仅需约五分钟即可适应每位患者的情况。该模型能在几秒钟内自动将一位患者的X射线与3D扫描结果相匹配,精度达到亚毫米级。
Named xvr (which stands for X-ray volume registration), it outperformed existing AI methods by an order of magnitude across a wide range of patients, body parts, and medical procedures.
该系统名为xvr(代表X-ray volume registration,即X射线体积配准),在广泛的各类患者、身体部位和医疗程序中,其性能比现有AI方法高出数量级。
“A majority of Americans live more than an hour away from a center that can perform noninvasive procedures, like emergency stroke interventions. An hour in stroke time is incredibly substantial. Making these procedures easier by combining 2D and 3D information enables these types of highly specialized life-saving procedures to be more accessible to much broader parts of the population,” says Vivek Gopalakrishnan, a postdoc in the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL); a recent graduate of the Harvard-MIT Program in Health Sciences and Technology; and lead author of a paper on xvr, which appears today in Nature.
“大多数美国人居住在距离能执行非侵入性程序(如急诊卒中干预)的医疗中心一小时以上的地方。在卒中救治中,一小时的时间至关重要。通过将2D和3D信息相结合,使这些程序更加简便,能让这类高度专业化的救命措施惠及更广泛的人群。”Vivek Gopalakrishnan表示,他是麻省理工学院计算机科学与人工智能实验室(CSAIL)的博士后研究员;哈佛-麻省理工学院健康科学与技术项目的近期毕业生;也是今天发表在《自然》杂志上关于xvr论文的主要作者。
He is joined on the paper by his advisor Polina Golland, the Sunlin and Priscilla Chou Professor of Electrical Engineering and Computer Science (EECS), a principal investigator in CSAIL, the leader of the Medical Vision Group, and co-senior author of the paper; and Neel Dey, a former postdoc in the Medical Vision Group who is now an investigator at Harvard Medical School and Massachusetts General Hospital as well as co-senior author on the paper. Additional co-authors include David-Dimitris Chlorogiannis, a researcher and clinician at Harvard Medical School; Andrew Abumoussa, a neurosurgeon at St. Luke’s Marion Bloch Neuroscience Institute; Anna M. Larson, a pediatric clinician at Shriners Children’s Hospital; Nazim Haouchine, an assistant professor of radiology at Harvard and Brigham and Women’s Hospital; Darren B. Orbach, a physician and scientist at Boston Children’s Hospital; and Sarah Frisken, an associate professor of radiology at Harvard.
论文的其他作者包括他的导师 Polina Golland,她是电气工程与计算机科学(EECS)的 Sunlin 和 Priscilla Chou 教授、CSAIL 的首席研究员、医学视觉小组的负责人以及论文的联合通讯作者;以及 Neel Dey,他是医学视觉小组的前博士后,现任哈佛医学院和马萨诸塞州总医院的研究员,同时也是论文的联合通讯作者。其他合著者还包括:David-Dimitris Chlorogiannis,哈佛医学院的研究人员和临床医生;Andrew Abumoussa,圣卢斯·马林·布洛赫神经科学研究所的神经外科医生;Anna M. Larson,Shriners 儿童医院的小儿临床医生;Nazim Haouchine,哈佛大学和布里格姆妇女医院的放射学助理教授;Darren B. Orbach,波士顿儿童医院的医师兼科学家;以及 Sarah Frisken,哈佛大学的放射学副教授。
Making X-rays more informative
使 X 射线提供更多有用信息
In many minimally invasive surgical procedures, like angioplasty to open blocked arteries, clinicians insert instruments through a tiny incision and use a high-speed mobile X-ray scanner to generate images that allow them to visualize the procedure from any angle.
在许多微创手术中,例如用于疏通阻塞动脉的血管成形术,临床医生会通过微小的切口插入器械,并使用高速移动 X 射线扫描仪生成图像,从而能够从任何角度观察手术过程。
But to guide surgical tools without accidentally damaging other tissue, clinicians must align real-time X-rays with the patient’s preoperative MRI or CT scan. This process, called registration, helps them determine where the tool is in relation to anatomical structures.
但为了在不意外损伤其他组织的情况下引导手术工具,临床医生必须将实时 X 射线图像与患者的术前 MRI 或 CT 扫描结果进行对齐。这一称为配准的过程有助于他们确定工具相对于解剖结构的位置。
“It takes decades of training for a clinician to become skilled enough to see grainy, 2D images and understand how everything is oriented. We want to make these 2D X-rays more informative, so it becomes safer and easier to do these life-saving procedures,” Gopalakrishnan says.
"临床医生需要数十年的训练才能熟练地解读颗粒状的二维图像并理解各部分的方位关系。我们希望使这些二维 X 射线图像包含更多信息,从而使这些挽救生命的手术更安全、更容易操作," Gopalakrishnan 表示。
Manual registration methods are slow and burdensome, requiring the clinician to guess the position of a surgical instrument by punching numbers into a computer or clicking anatomical landmarks on a screen.
手动配准方法缓慢且繁琐,要求临床医生通过向计算机输入数字或在屏幕上点击解剖标志点来猜测手术器械的位置。
To streamline the process, researchers are developing AI models that can predict 2D/3D registration. But people have such diverse anatomy that a model which works well for some patients may fail for others.
为了简化这一流程,研究人员正在开发能够预测 2D/3D 配准的 AI 模型。但由于人体解剖结构差异巨大,适用于某些患者的模型可能在其他患者身上失效。
A lack of high-quality annotated medical image data makes it difficult to train a deep-learning model robust enough to adapt to many patients, Gopalakrishnan says.
Gopalakrishnan 表示,缺乏高质量标注的医学图像数据使得难以训练出足够鲁棒以适应众多患者的深度学习模型。
Rather than trying to make a machine-learning model that can be applied to all patients, the researchers built a model designed to adapt extremely well for the specific patient.
研究人员没有试图构建一个适用于所有患者的机器学习模型,而是构建了一个专为特定患者量身定制、适应性极强的模型。
“We tailor this one specific model for this one specific patient, and it doesn’t matter if it works on other people because there will be different models for those people,” Gopalakrishnan adds.
“我们为每一位特定患者定制这一个特定模型,它是否对其他人生效并不重要,因为那些人会有不同的模型。”Gopalakrishnan 补充道。
Patient-specific machine learning
患者特异性机器学习
Xvr takes one patient’s preoperative 3D scan, like an MRI or CT, and uses it to generate thousands of synthetic X-rays from many angles, producing about 1,000 images each second. It uses a physics-based simulation of the X-ray process to ensure these synthetic images are realistic.
Xvr 获取一位患者的术前 3D 扫描数据(如 MRI 或 CT),并利用这些数据从多个角度生成数千张合成 X 射线图像,每秒可生成约 1,000 张图像。它采用基于物理的 X 射线过程模拟,以确保这些合成图像具有真实感。
“Instead of generating data from nothing, like some types of generative AI, this physics simulation is entirely based on the CT scan or MRI from this patient. Because xvr creates patient-specific data in a purely physics-based manner, there is no room for hallucinations,” Gopalakrishnan says.
“与某些类型的生成式 AI 从零开始生成数据不同,这种物理模拟完全基于该患者的 CT 扫描或 MRI。由于 xvr 以纯物理方式创建患者特异性数据,因此不存在幻觉问题,”Gopalakrishnan 说。
The xvr framework uses these simulated data to train an AI model that can accurately align this patient’s 2D X-rays with their 3D image scan in a matter of seconds.
xvr 框架利用这些模拟数据来训练一个 AI 模型,该模型能够在几秒钟内将该患者的 2D X 射线与其 3D 图像扫描准确对齐。
But while such a registration model is highly accurate, it would take about 12 hours to train from scratch for each patient, making it impossible to deploy in an emergency. To make the process faster, the researchers used xvr to pretrain a more versatile AI system, called a foundation model, that can quickly adjust to each new patient.
然而,尽管此类配准模型高度准确,但为每位患者从头训练需要约 12 小时,这使得在紧急情况下部署变得不可能。为了加快这一过程,研究人员使用 xvr 预训练了一个更通用的 AI 系统——基础模型(foundation model),该系统可以快速适应每位新患者。
They collected whole-body 3D medical scans from more than 2,000 patients covering a wide range of ages, image modalities, and regions. Xvr used these diverse data to generate synthetic X-rays and train a foundation model to perform 2D/3D registration.
他们收集了来自 2,000 多名患者的全身 3D 医学扫描数据,涵盖了广泛的年龄、成像模态和身体区域。Xvr 利用这些多样化数据生成合成 X 射线,并训练基础模型执行 2D/3D 配准。
This pretrained model can adapt to a new patient in about five minutes, and performs registration with the same accuracy as if it had been trained from scratch.
这个预训练模型可以在大约五分钟内适应新患者,并且其配准精度与从头训练的效果相同。
“So now you can get patient-specific accuracy but also in a very rapid time frame,” Gopalakrishnan says.
“因此,你现在既能获得患者特异性的精度,又能在非常短的时间内完成,”Gopalakrishnan 说。
The team tested the model on the largest available dataset of real 2D/3D registrations, incorporating data from five hospitals that covered dozens of bones and organ systems in adult and pediatric patients.
团队在最大的可用真实 2D/3D 配准数据集上对该模型进行了测试,其中包含了来自五家医院的数据,涵盖了成人和儿科患者的数十种骨骼和器官系统。
Xvr significantly outperformed other AI-based methods in accuracy and robustness, while operating fast enough for emergency surgeries. The model could also be used to improve the performance of robotic surgery technologies.
Xvr 在准确性和鲁棒性方面显著优于其他基于 AI 的方法,同时运行速度足够快,可满足急诊手术的需求。该模型还可用于提升机器人手术技术的性能。
In the future, the researchers hope to focus on making xvr faster for real-time deployment, conducting further studies to verify its reliability in additional situations, and extending the system to handle more complex scenarios, like moving body parts.
未来,研究人员希望专注于使 xvr 更快以实现实时部署,开展进一步研究以验证其在更多场景下的可靠性,并将系统扩展至处理更复杂的场景,例如移动身体部位。
“For the past two years, we’ve been carefully developing this algorithm and validating it. Now, we are collaborating closely with surgical robotics companies and clinical groups to turn this research into useful tools for navigation or deployment,” Gopalakrishnan says.
“在过去两年里,我们一直在仔细开发并验证该算法。现在,我们正在与手术机器人公司和临床团队紧密合作,将这项研究转化为用于导航或部署的实用工具,”Gopalakrishnan 表示。
This work was funded, in part, but the National Institutes of Health (NIH), the MIT CSAIL-Wistron Program, the MIT-IBM Computing Research Lab, the MIT Jameel Clinic, the MIT Health and Life Sciences Collaborative, and the Chou Family Transformative Research Fund.
这项工作部分获得了美国国立卫生研究院(NIH)、MIT CSAIL-Wistron 计划、MIT-IBM 计算研究实验室、MIT Jameel 诊所、MIT 健康与生命科学合作项目以及 Chou 家族变革性研究基金的资助。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力