哈佛新论文:生成模型或缺失第三缩放轴
New Harvard paper shows generative models may be missing a third scaling axis: h…
New Harvard paper shows generative models may be missing a third scaling axis: how much they explore during training.
哈佛大学新论文显示,生成模型可能缺失第三个扩展轴:训练期间探索的程度。
What if the next scaling axis for generative models is not a bigger model or more data, but more candidate generations per training step?
如果生成模型的下一扩展轴不是更大的模型或更多的数据,而是每个训练步骤中更多的候选生成呢?
Added to a strong Representation Autoencoder (RAE) image-generation recipe, exploration reaches the baseline's final performance with 6.2× fewer training samples processed and 4.1× fewer FLOPs.
在强大的表示自编码器(RAE)图像生成方案中加入探索,以6.2倍更少的训练样本处理和4.1倍更少的FLOPs达到了基线的最终性能。
This paper treats best-of-K training as something bigger: a way to scale how many modes a generative model can learn without adding inference steps.
这篇论文将最佳K训练视为更宏大的概念:一种在不增加推理步骤的情况下扩展生成模型可学习模式数量的方式。
Today, diffusion, flow, and autoregressive models handle multimodal targets largely by splitting generation into many easier steps, which also creates a mismatch between how they train and how they sample.
如今,扩散模型、流模型和自回归模型主要通过将生成拆分为许多更简单的步骤来处理多模态目标,这也造成了训练与采样之间的不匹配。
Explorative Modeling moves that burden into training instead.
探索性建模将这一负担转移到训练阶段。
At each update, the model tries K candidate generations and learns only from the candidate closest to the target, letting different latents specialize to different modes rather than being pulled toward an average.
在每次更新中,模型尝试K个候选生成,并仅从最接近目标的候选中学习,使不同的潜在变量专门化于不同模式,而不是被拉向平均值。
If the scaling trend survives larger runs, compute-optimal generative training may need to budget for exploration alongside parameters and data.
如果这种扩展趋势在更大规模的运行中持续,计算最优的生成训练可能需要在参数和数据之外,为探索分配预算。
– arxiv. org/abs/2607.27372
– arxiv.org/abs/2607.27372
Title: "Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation"
标题:“探索性建模:解锁第三个预训练轴与端到端生成”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力