跳到主内容
@wquguru
精选80Rohan Paul模型发布/更新多源精选 ×4

智谱唐杰谈AI扩展范式:参数增长不再是唯一标尺

Brilliant piece by Zhipu Founder Tang Jie.

原文
发到 X

Brilliant piece by Zhipu Founder Tang Jie.

AI scaling is moving past parameter growth.

“How many parameters?” is becoming a weak way to describe how capable a model should be.

That model scaling now has several independent dials: parameters, training data, compute per forward pass, and post-training.

The best place to spend the next unit of compute depends on what the model needs to do and how often it will be used.

“Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training.”

“Put inference into the objective and the optimum moves toward smaller models trained far longer.”

GLM-5.3 is Tang’s implented example: same base, architecture, total parameters, and activated parameters as GLM-5.2, but 1 more month spent on long-horizon environments and RL.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →