跳到主内容
精选85Rohan Paul论文研究

Meta FAIR新论文:Chinchilla缩放定律存在盲点

A new Meta FAIR paper finds a blind spot in the Chinchilla scaling law that beco…

原文
推荐理由

做模型缩放和训练成本优化的同学必看,Skaling 用一个小改动大幅提升外推精度,建议精读论文并复现对比。

A new Meta FAIR paper finds a blind spot in the Chinchilla scaling law that becomes expensive when you extrapolate.

Meta FAIR 的一篇新论文发现,Chinchilla 缩放定律中存在一个盲点,当外推时,这个盲点会变得代价高昂。

Chinchilla can look almost perfect inside a training grid and still mispredict what happens at the frontier.

Chinchilla 在训练网格内可能看起来几乎完美,但仍然会在前沿预测上出错。

The problem is: Chinchilla assumes model size and training data help separately, but the experiments show that each changes how useful the other one is.

问题在于:Chinchilla 假设模型大小和训练数据是独立起作用的,但实验表明,它们各自会改变对方的有用程度。

Skaling adds just 1 extra term to capture that connection, cutting prediction error by about 1.5–3× and getting full-grid Chinchilla-level prediction accuracy with roughly 10× less profiling compute.

Skaling 仅增加了一个额外项来捕捉这种关联,将预测误差降低了约 1.5–3 倍,并且用大约 10 倍少的性能分析计算量,就达到了全网格 Chinchilla 级别的预测精度。

On Farseer, the difference becomes huge at frontier scale: at 2×10^25 FLOPs, Chinchilla points to ~380 tokens per parameter, while Skaling and the paper's direct estimates land around 20–40.

在 Farseer 上,这种差异在前沿规模下变得巨大:在 2×10^25 FLOPs 时,Chinchilla 指向约 380 个 token 每参数,而 Skaling 和论文的直接估计落在 20–40 左右。

– arxiv. org/abs/2608.07222

– arxiv.org/abs/2608.07222

Title: "Skaling: Chinchilla's Exponents Meet Kaplan's Coupling"

标题:“Skaling:Chinchilla 的指数遇上 Kaplan 的耦合”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近