跳到主内容
@wquguru
精选75elvis论文研究

Meta提出Skaling law:耦合容量与数据的缩放定律

Impressive new paper from Meta.

原文
发到 X

Impressive new paper from Meta.

Meta 令人印象深刻的新论文。

(bookmark it)

(收藏它)

Scaling laws assume model size and training data act on loss independently.

缩放定律假设模型大小和训练数据对损失的影响是独立的。

This work introduces Skaling law, which couples capacity and data through a single interaction exponent. The extra term cuts mean absolute percentage error by 1.5x to 3x across both interpolation and extrapolation.

这项工作引入了 Skaling 定律,通过一个交互指数将容量和数据耦合起来。额外的项在插值和外推中均将平均绝对百分比误差降低了 1.5 倍到 3 倍。

The largest corrections land in the data-scarce and heavy-overtraining regimes where the standard Chinchilla and Kaplan forms drift.

最大的修正出现在数据稀缺和重度过训练区域,而标准的 Chinchilla 和 Kaplan 形式在这些区域会偏离。

Paired with a sparse grid restricted to low-compute runs, it extrapolates the full grid using roughly 10x less compute than a uniform sweep.

与限制在低计算量运行的稀疏网格相结合,它使用比均匀扫描少约 10 倍的计算量来外推整个网格。

Why does it matter?

为什么这很重要?

Deployment now happens well past compute optimal. A law that stays accurate there, and that can be fit from small runs, changes how a pretraining budget gets planned.

现在部署远在计算最优之后。一个在该区域保持准确且能从少量运行中拟合的定律,会改变预训练预算的规划方式。

Paper: https://arxiv.org/abs/2608.07222

论文:https://arxiv.org/abs/2608.07222

Track more trending AI papers in our academy: https://academy.dair.ai/

在我们的学院中追踪更多热门 AI 论文:https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近