Meta 新论文:小模型预测规模效应不佳或因欠调参
New Meta paper shows, small models may not be bad predictors of scale; they may…
New Meta paper shows, small models may not be bad predictors of scale; they may just be getting under-tuned.
Meta 的新论文表明,小模型可能并非规模的不良预测器;它们可能只是调优不足。
Finds scaling laws emerge around 4M parameters, where models can train in under 1 hour on 1 GPU.
研究发现,缩放定律在约 4M 参数时出现,此时模型可在 1 块 GPU 上于 1 小时内完成训练。
Small models are unusually sensitive to hyperparameters.
小模型对超参数异常敏感。
With 4 or 16 configurations per scale, the law is basically invisible; at 64 it appears but extrapolates poorly, and at 256 it becomes accurate.
每个规模只有 4 或 16 种配置时,该定律基本不可见;64 种时它出现但外推效果不佳;256 种时则变得准确。
As models get larger, good settings occupy more of the search space, while the effective number of hyperparameters near the optimum drops toward 1.
随着模型变大,好的设置占据更多搜索空间,而最优值附近的超参数有效数量降至接近 1。
That helps explain why scaling laws look cleaner at larger sizes: the models are easier to tune.
这有助于解释为什么缩放定律在更大规模时看起来更清晰:模型更容易调优。
As a check, small-scale runs recover that pre-norm transformers scale better than post-norm over the tested range.
作为验证,小规模运行重现了在测试范围内 pre-norm transformer 的缩放效果优于 post-norm 的结果。
There is a limit: extrapolate too far beyond the measured scales, and statistical errors can dominate.
存在一个限制:外推超出实测规模太远时,统计误差可能占主导。
For model research, cheap experiments may need more tuning breadth, not more model size.
对于模型研究,廉价实验可能需要更广的调优范围,而非更大的模型规模。
– arxiv. org/abs/2608.11859
– arxiv. org/abs/2608.11859
Title: "Small-Scale Experiments: Are They There Yet?"
标题:“小规模实验:它们已经准备好了吗?”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力