跳到主内容
精选80Rohan Paul模型发布/更新

智谱GLM-5.3漏洞发现超越Anthropic Mythos 5

China's Z .ai;s new model GLM-5.3 beats Anthropic's Mythos 5 in cyber-defence te…

原文
推荐理由

GLM-5.3在漏洞发现上超越Anthropic旗舰,编码能力大幅跃升,做安全或模型评估的同学值得关注,建议跟进权重发布后的实测。

China's Z .ai;s new model GLM-5.3 beats Anthropic's Mythos 5 in cyber-defence tests, i.e. at vulnerability discovery.

中国的Z .ai的新模型GLM-5.3在网络安全防御测试中击败了Anthropic的Mythos 5,即在漏洞发现方面。

Coding jumped hard: Terminal Bench 3.0 went 4.6 → 28.3 and DeepSWE 46.2 → 66.9.

编码能力大幅提升:Terminal Bench 3.0从4.6跃升至28.3,DeepSWE从46.2升至66.9。

On CyberGym, which tests finding and validating source-code flaws, GLM-5.3 reports 84.5% vs. 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol.

在测试发现和验证源代码缺陷的CyberGym上,GLM-5.3报告84.5%的成绩,而Mythos 5为83.8%,GPT-5.6 Sol为83.6%。

But on ExploitBench, which requires deeper exploit development, GLM-5.3 drops to 54.4%, vs. 78.0% for Mythos 5 and 76.5% for GPT-5.6 Sol.

但在需要更深入漏洞利用开发的ExploitBench上,GLM-5.3降至54.4%,而Mythos 5为78.0%,GPT-5.6 Sol为76.5%。

The same gap appears on ExploitGym, where Z .ai reports 105 completed tasks in 2 hours and 130 in 6, against Mythos 5's 181 and 247.

同样的差距出现在ExploitGym上,Z .ai报告在2小时内完成105项任务,6小时内完成130项,而Mythos 5分别为181项和247项。

Its training environments increasingly resemble multi-day engineering work, forcing the model to diagnose, edit, test, and recover across long task chains.

其训练环境越来越类似于多日的工程工作,迫使模型在长任务链中进行诊断、编辑、测试和恢复。

Z .ai plans to publish GLM-5.3's weights after a two-week safety review, while limiting its most sensitive cyber functions to verified users.

Z .ai计划在两周安全审查后发布GLM-5.3的权重,同时将其最敏感的网络安全功能限制给经过验证的用户。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近