Meta FAIR新论文:Chinchilla缩放定律存在盲点
A new Meta FAIR paper finds a blind spot in the Chinchilla scaling law that beco…
做模型缩放和训练成本优化的同学必看,Skaling 用一个小改动大幅提升外推精度,建议精读论文并复现对比。
A new Meta FAIR paper finds a blind spot in the Chinchilla scaling law that becomes expensive when you extrapolate.
Meta FAIR 的一篇新论文发现,Chinchilla 缩放定律中存在一个盲点,当外推时,这个盲点会变得代价高昂。
Chinchilla can look almost perfect inside a training grid and still mispredict what happens at the frontier.
Chinchilla 在训练网格内可能看起来几乎完美,但仍然会在前沿预测上出错。
The problem is: Chinchilla assumes model size and training data help separately, but the experiments show that each changes how useful the other one is.
问题在于:Chinchilla 假设模型大小和训练数据是独立起作用的,但实验表明,它们各自会改变对方的有用程度。
Skaling adds just 1 extra term to capture that connection, cutting prediction error by about 1.5–3× and getting full-grid Chinchilla-level prediction accuracy with roughly 10× less profiling compute.
Skaling 仅增加了一个额外项来捕捉这种关联,将预测误差降低了约 1.5–3 倍,并且用大约 10 倍少的性能分析计算量,就达到了全网格 Chinchilla 级别的预测精度。
On Farseer, the difference becomes huge at frontier scale: at 2×10^25 FLOPs, Chinchilla points to ~380 tokens per parameter, while Skaling and the paper's direct estimates land around 20–40.
在 Farseer 上,这种差异在前沿规模下变得巨大:在 2×10^25 FLOPs 时,Chinchilla 指向约 380 个 token 每参数,而 Skaling 和论文的直接估计落在 20–40 左右。
– arxiv. org/abs/2608.07222
– arxiv.org/abs/2608.07222
Title: "Skaling: Chinchilla's Exponents Meet Kaplan's Coupling"
标题:“Skaling:Chinchilla 的指数遇上 Kaplan 的耦合”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力