跳到主内容
精选88Microsoft Research(RSS)模型发布/更新

微软发布GigaPath-Flash与GigaTIME-Flash高效病理大模型

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

原文
推荐理由

病理AI领域的重磅轻量化更新,用极小参数实现接近原版SOTA的性能,大幅降低算力门槛,适合做医疗影像研究的团队直接复用。

At a glance

概览

  • The Flash family extends GigaPath and GigaTIME with dramatically improved efficiency, making large-scale pathology research more accessible and practical.
  • A distilled pathology foundation model backbone reduces computational requirements without sacrificing performance, enabling repeated analyses across larger patient cohorts.
  • These open models support population-scale discovery, helping researchers investigate disease biology, biomarkers, and clinical outcomes across diverse cancer datasets.
  • Flash 系列在大幅提升效率的基础上扩展了 GigaPath 和 GigaTIME,使大规模病理学研究更加普及且实用。
  • 蒸馏后的病理基础模型骨干网络在不牺牲性能的前提下降低了计算需求,使得针对更大患者队列的重复分析成为可能。
  • 这些开源模型支持人群规模的发现,帮助研究人员在各种癌症数据集中探究疾病生物学、生物标志物和临床结果。

GigaPath (opens in new tab) and GigaTIME (opens in new tab) demonstrated how foundation models can support whole-slide analysis and tumor microenvironment modeling from routinely collected pathology data. GigaPath-Flash and GigaTIME-Flash make these capabilities substantially more efficient, enabling researchers to analyze larger cohorts, run more experiments, and move toward population-scale discovery. GigaPath-Flash and GigaTIME-Flash are research models. They are not intended or validated for clinical use, including diagnosis, prognosis, treatment selection, or other patient-care decisions. Performance may vary across datasets, scanners, institutions, populations, and use cases.

GigaPath(新标签页打开)和 GigaTIME(新标签页打开)展示了基础模型如何支持从常规收集的病理数据进行全切片分析和肿瘤微环境建模。GigaPath-Flash 和 GigaTIME-Flash 使这些能力显著更高效,使研究人员能够分析更大的队列、运行更多的实验,并迈向人群规模的发现。GigaPath-Flash 和 GigaTIME-Flash 是研究模型。它们并非旨在或验证用于临床用途,包括诊断、预后、治疗选择或其他患者护理决策。性能在不同数据集、扫描仪、机构、人群和使用场景中可能会有所不同。

The scale opportunity in computational pathology

计算病理学中的规模机遇

Histopathology is among the richest and most widely available sources of information in cancer research. Every tissue biopsy produces a whole-slide image that captures cellular morphology at subcellular resolution — and hospitals generate millions of these slides each year. This data contains information relevant to diagnosis, prognosis, treatment selection, and the biology of the tumor microenvironment.

组织病理学是癌症研究中信息最丰富且最广泛可用的来源之一。每次组织活检都会生成一张全切片图像,以亚细胞分辨率捕捉细胞形态——医院每年生成数百万张这样的切片。这些数据包含与诊断、预后、治疗选择和肿瘤微环境生物学相关的信息。

Foundation models have begun to unlock this information at scale. But whole-slide images are large — often exceeding a gigapixel — and applying a foundation model to even a single slide requires processing thousands of image tiles. When a research question involves tens of thousands of patients, the computational cost grows quickly. And population-scale discovery is not a single model run: it requires repeated cycles of feature extraction, statistical analysis, hypothesis testing, and validation across patient subgroups, biomarkers, and clinical endpoints.

基础模型已开始大规模解锁这些信息。但全切片图像很大——通常超过十亿像素——将基础模型应用于单个切片就需要处理数千个图像图块。当研究问题涉及数万名患者时,计算成本会迅速增长。而且,人群规模的发现并非单次模型运行:它需要在患者亚组、生物标志物和临床终点之间进行特征提取、统计分析、假设检验和验证的重复循环。

Computational cost limits the number of patients, datasets, tasks, and hypotheses that researchers can study. To realize the full potential of pathology foundation models, we need models that can be applied repeatedly and affordably across large patient populations.

计算成本限制了研究人员可以探索的患者数量、数据集、任务和假设。为了充分发挥病理基础模型的潜力,我们需要能够在大型患者群体中反复且经济地应用的模型。

From GigaPath and GigaTIME to the -Flash family

从 GigaPath 和 GigaTIME 到 Flash 家族

GigaPath (Nature, 2024) (opens in new tab) is a whole-slide foundation model pretrained on large-scale real-world histopathology data from Providence. Unlike models that operate only at the tile level, GigaPath learns contextualized representations of entire slides, capturing both local cellular patterns and global tissue architecture.

GigaPath(《自然》,2024)(在新标签页中打开)是一个全切片基础模型,在来自 Providence 的大规模真实世界组织病理学数据上进行预训练。与仅在图块级别运行的模型不同,GigaPath 学习整个切片的上下文表示,捕捉局部细胞模式和全局组织结构。

GigaTIME (Cell, 2026) (opens in new tab) extends this line of work to tumor microenvironments. Trained on 40 million cells with paired H&E and multiplex immunofluorescence (mIF) data, GigaTIME translates routine H&E images into virtual spatial proteomics maps across 21 protein channels. Applied to over 14,000 cancer patients, it generated a virtual population that uncovered more than 1,200 statistically significant associations between immune cell states and clinical biomarkers.

GigaTIME(《细胞》,2026)(在新标签页中打开)将这一研究方向扩展至肿瘤微环境。GigaTIME 使用带有配对 H&E 和多色免疫荧光 (mIF) 数据的 4000 万个细胞进行训练,能够将常规 H&E 图像转化为跨越 21 个蛋白通道的虚拟空间蛋白质组学图谱。应用于超过 14,000 名癌症患者后,它生成了一个虚拟人群,揭示了免疫细胞状态与临床生物标志物之间超过 1,200 个具有统计学显著性的关联。

GigaPath and GigaTIME addressed the scale of pathology data and biological discovery. And now, the Flash family of models addresses scale of experimentation.

GigaPath 和 GigaTIME 解决了病理数据的规模和生物学发现的问题。而现在,Flash 系列模型则致力于解决实验的规模问题。

Figure 1: Overview of the GigaPath/GigaTIME model family. GigaPath-Flash provides efficient tile and slide encoders at 22M parameters. GigaTIME-Flash predicts spatial proteomics from H&E, replacing the CNN backbone with the distilled ViT-S encoder.

图 1:GigaPath/GigaTIME 模型家族概览。GigaPath-Flash 提供拥有 2200 万参数的图块和切片编码器。GigaTIME-Flash 从 H&E 预测空间蛋白质组学,用蒸馏后的 ViT-S 编码器替换了 CNN 主干网络。

Introducing GigaPath-Flash and GigaTIME-Flash

介绍 GigaPath-Flash 和 GigaTIME-Flash

The Flash family shares a core design goal: preserving useful pathology representations while substantially reducing the computational resources required to generate and use them. Both models are built on a common efficient backbone — a compact ViT-S tile encoder distilled from the original billion-parameter GigaPath encoder — and both are released under the Apache 2.0 license.

Flash 系列共享一个核心设计目标:在大幅降低生成和使用这些表示所需计算资源的同时,保留有用的病理学表示。这两个模型都构建在一个共同的高效主干网络上——一个由原始十亿参数 GigaPath 编码器蒸馏而来的紧凑 ViT-S 图块编码器——并且都以 Apache 2.0 许可证发布。

GigaPath-Flash

GigaPath-Flash

GigaPath-Flash is an efficient foundation model for whole-slide representation learning. It combines a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder. The tile encoder is distilled from the original GigaPath ViT-g teacher, transferring the representational capacity of a billion-parameter model into a backbone that is an order of magnitude smaller. The slide encoder contextualizes all tile embeddings via dilated attention, scaling linearly with the number of tiles.

GigaPath-Flash 是一个用于全切片表示学习的高效基础模型。它将一个拥有 2200 万参数的 ViT-S 图块编码器与一个拥有 2100 万参数的 LongNet 切片编码器相结合。该图块编码器是从原始的 GigaPath ViT-g 教师模型蒸馏而来,将十亿参数模型的表示能力转移到了一个规模小一个数量级的主干网络中。切片编码器通过空洞注意力机制对所有图块嵌入进行上下文化处理,其复杂度随图块数量线性扩展。

On slide-level classification benchmarks (PANDA prostate grading and EBRAINS brain tumor subtyping), GigaPath-Flash achieves the lowest inference cost among whole-slide pretrained models while retaining competitive performance — scoring within 3% of the original GigaPath at roughly 50 times less compute.

在切片级分类基准测试(PANDA前列腺分级和EBRAINS脑肿瘤亚型分类)中,GigaPath-Flash 在所有全切片预训练模型中实现了最低的推理成本,同时保持了具有竞争力的性能——其得分仅比原始 GigaPath 低 3%,而计算量仅为后者的约 1/50。

Figure 2: Efficiency–performance trade-off on whole-slide benchmarks. GigaPath-Flash (red, top-left) achieves competitive performance at substantially lower computational cost than other whole-slide pretrained models.

图 2:全切片基准测试上的效率与性能权衡。GigaPath-Flash(红色,左上角)以显著更低的计算成本实现了与其他全切片预训练模型相媲美的性能。

GigaTIME-Flash

GigaTIME-Flash

GigaTIME-Flash replaces the CNN backbone of the original GigaTIME with the GigaPath-Flash ViT-S encoder, paired with a lightweight convolutional decoder for H&E-to-mIF translation. The model is fine-tuned using LoRA adapters, keeping the pretrained encoder weights largely frozen.

GigaTIME-Flash 使用 GigaPath-Flash ViT-S 编码器替换了原始 GigaTIME 的 CNN 主干网络,并搭配一个轻量级卷积解码器用于 H&E 到 mIF 的转换。该模型使用 LoRA 适配器进行微调,使预训练编码器的权重基本保持冻结状态。

On both in-distribution and out-of-distribution cohorts spanning brain, breast, colon, and lung cancers, GigaTIME-Flash matches or improves upon the original GigaTIME in spatial protein prediction quality. The gains are particularly notable on out-of-distribution data, suggesting that the foundation model backbone improves generalization to previously unseen tissue types.

在涵盖脑癌、乳腺癌、结肠癌和肺癌的分布内和分布外队列中,GigaTIME-Flash 的空间蛋白预测质量匹配或优于原始 GigaTIME。这种提升在分布外数据上尤为显著,表明基础模型主干提升了泛化能力,使其能够适应此前未见过的组织类型。

Figure 3: Mean windowed Pearson correlation for GigaTIME and GigaTIME-Flash on the GigaTIME test set and four out-of-distribution Prov-TMA cohorts. GigaTIME-Flash matches or improves upon the original across all cohorts.

图 3:GigaTIME 和 GigaTIME-Flash 在 GigaTIME 测试集及四个分布外 Prov-TMA 队列上的平均窗口皮尔逊相关系数。GigaTIME-Flash 在所有队列中均匹配或优于原始模型。

Efficiency without giving up the foundation

无需牺牲基础模型特性的效率提升

The efficiency gains of the Flash models are substantial:

Flash 模型的效率提升显著:

ModelTypeEfficiency gain
GigaPath-FlashWhole-slide Foundation Model~50× less compute, 97% of predictive performance compared to GigaPath
GigaTIME-FlashSpatial Proteomics~6× faster, ~8× less memory, better predictive performance compared to GigaTIME
模型类型效率增益
GigaPath-Flash全切片基础模型计算量减少约 50 倍,预测性能达到 GigaPath 的 97%
GigaTIME-Flash空间蛋白质组学速度提升约 6 倍,内存占用减少约 8 倍,预测性能优于 GigaTIME

For a single slide, these differences reduce runtime and hardware requirements. Across tens of thousands of slides, they can determine whether an experiment is practical at all. To illustrate, we estimate the wall-clock time for generating virtual mIF across cohorts of different sizes on a single A100 GPU, assuming approximately 10,000 tiles per slide:

对于单个切片,这些差异减少了运行时间和硬件需求。而在处理数万张切片时,它们可能决定实验是否具备实际可行性。为了说明这一点,我们估算了在单张 A100 GPU 上为不同规模的队列生成虚拟 mIF 所需的墙钟时间,假设每张切片包含约 10,000 个图块:

Model1,000 slides100K slides1M slides
GigaTIME-Flash~2 GPU-hours~7 GPU-days~70 GPU-days
GigaTIME~7 GPU-hours~30 GPU-days~300 GPU-days
模型1,000 张切片10 万张切片100 万张切片
GigaTIME-Flash~2 GPU 小时~7 GPU 天~70 GPU 天
GigaTIME~7 GPU 小时~30 GPU 天~300 GPU 天

Estimates assume ~10,000 tiles per slide, batch size 128, single NVIDIA A100 GPU. Actual runtime depends on slide size, tiling resolution, and hardware.

估算基于每张切片约 10,000 个图块、批量大小 128、单张 NVIDIA A100 GPU。实际运行时间取决于切片尺寸、分块分辨率和硬件配置。

Figure 4: GigaTIME efficiency scaling. Left: throughput (tiles/sec) vs. batch size. Right: peak GPU memory (GB) vs. batch size. GigaTIME-Flash scales to over 1,600 tiles/sec while using a fraction of the memory.

图 4:GigaTIME 的效率扩展。左图:吞吐量(tiles/sec)与批次大小(batch size)的关系。右图:峰值 GPU 内存(GB)与批次大小的关系。GigaTIME-Flash 可扩展至超过 1,600 tiles/sec,同时仅使用极小比例的内存。

An open model release

开源模型发布

Both GigaPath-Flash and GigaTIME-Flash are released as open-weight models under the Apache 2.0 license. Model weights and code are available on HuggingFace:

GigaPath-Flash 和 GigaTIME-Flash 均以 Apache 2.0 许可证作为开放权重模型发布。模型权重和代码可在 HuggingFace 上获取:

  • GigaPath-Flash (opens in new tab)
  • GigaTIME-Flash (opens in new tab)
  • GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
  • GigaPath-Flash(在新标签页中打开)
  • GigaTIME-Flash(在新标签页中打开)
  • GigaPath-Flash 和 GigaTIME-Flash:用于全切片和肿瘤微环境分析的高效病理学基础模型

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近