跳到主内容
@wquguru
精选80Rohan Paul论文研究

Meta与牛津研究:多模态模型图像生成数据可大幅减少

The Meta/Oxford study finds, a multimodal model may need surprisingly little ima…

原文
发到 X

The Meta/Oxford study finds, a multimodal model may need surprisingly little image-generation data if language and visual understanding are trained with it from the start.

Meta/Oxford 的研究发现,如果从一开始就同时训练语言和视觉理解,多模态模型可能只需要极少量的图像生成数据。

So, you probably don’t need to spend that much training compute teaching a multimodal model to generate images.

所以,你可能不需要花费那么多训练计算量来教多模态模型生成图像。

Language training already helps the model with vision, and learning to understand images also makes it better at generating them. But training it to generate images does very little for language or image understanding.

语言训练已经帮助模型理解视觉,而学习理解图像也使其更擅长生成图像。但训练生成图像对语言或图像理解的提升微乎其微。

That leads to a very uneven training mix.

这导致了非常不均衡的训练配比。

In their 1T-token experiments, the best overall split was 70% language, 25% image understanding, and only 5% image generation.

在他们 1T token 的实验中,最佳的整体配比是 70% 语言、25% 图像理解,仅 5% 图像生成。

They then tested this at 13.5B scale over 2T tokens. Even with 5x fewer image-generation tokens than the balanced setup, GenEval improved from 0.467 to 0.482, while language and image understanding improved too.

他们随后在 13.5B 规模、2T token 上进行了测试。即使图像生成 token 比均衡设置少 5 倍,GenEval 从 0.467 提升到 0.482,同时语言和图像理解也有所提升。

There’s another lesson: don’t bolt vision on too late.

还有另一个教训:不要过晚地加入视觉能力。

The longer the model trains only on language, the more it starts ignoring the image and relying on language shortcuts. The authors call this “vision laziness.”

模型仅训练语言的时间越长,就越会开始忽略图像并依赖语言捷径。作者称之为“视觉懒惰”。

– arxiv. org/abs/2608.05000

– arxiv.org/abs/2608.05000

Title: "Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes"

标题:“迈向多模态预训练的物理:知识流动、模态协同、早期统一与配方”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近