跳到主内容
@wquguru
精选80Rohan Paul模型发布/更新

开源模型长周期网络能力差距缩至4-7个月

On long-horizon cyber capability, leading open-weight models now trail the close…

原文
发到 X

On long-horizon cyber capability, leading open-weight models now trail the closed frontier by only 4 to 7 months, down from 6 to 10 months through much of 2025.

GLM-5.2 matched closed models released about 4 months earlier. On long-horizon cyber ranges, GLM-5.2 matched Claude Opus 4.5, released roughly 7 months earlier.

And also that long-horizon cyber capability can now be scaled with compute.

On the AI Security Institute’s 32-step “The Last Ones” range, GPT-5.6 Sol completed the full 32-step range in 7 of 10 attempts.

With a 100M-token budget for each run.

Performance continued improving as the model received more inference tokens. i.e. operators can gain materially stronger cyber capability simply by spending more runtime compute, without retraining the model.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近