智谱GLM-5.3漏洞发现超越Anthropic Mythos 5
China's Z .ai;s new model GLM-5.3 beats Anthropic's Mythos 5 in cyber-defence te…
GLM-5.3在漏洞发现上超越Anthropic旗舰,编码能力大幅跃升,做安全或模型评估的同学值得关注,建议跟进权重发布后的实测。
China's Z .ai;s new model GLM-5.3 beats Anthropic's Mythos 5 in cyber-defence tests, i.e. at vulnerability discovery.
中国的Z .ai的新模型GLM-5.3在网络安全防御测试中击败了Anthropic的Mythos 5,即在漏洞发现方面。
Coding jumped hard: Terminal Bench 3.0 went 4.6 → 28.3 and DeepSWE 46.2 → 66.9.
编码能力大幅提升:Terminal Bench 3.0从4.6跃升至28.3,DeepSWE从46.2升至66.9。
On CyberGym, which tests finding and validating source-code flaws, GLM-5.3 reports 84.5% vs. 83.8% for Mythos 5 and 83.6% for GPT-5.6 Sol.
在测试发现和验证源代码缺陷的CyberGym上,GLM-5.3报告84.5%的成绩,而Mythos 5为83.8%,GPT-5.6 Sol为83.6%。
But on ExploitBench, which requires deeper exploit development, GLM-5.3 drops to 54.4%, vs. 78.0% for Mythos 5 and 76.5% for GPT-5.6 Sol.
但在需要更深入漏洞利用开发的ExploitBench上,GLM-5.3降至54.4%,而Mythos 5为78.0%,GPT-5.6 Sol为76.5%。
The same gap appears on ExploitGym, where Z .ai reports 105 completed tasks in 2 hours and 130 in 6, against Mythos 5's 181 and 247.
同样的差距出现在ExploitGym上,Z .ai报告在2小时内完成105项任务,6小时内完成130项,而Mythos 5分别为181项和247项。
Its training environments increasingly resemble multi-day engineering work, forcing the model to diagnose, edit, test, and recover across long task chains.
其训练环境越来越类似于多日的工程工作,迫使模型在长任务链中进行诊断、编辑、测试和恢复。
Z .ai plans to publish GLM-5.3's weights after a two-week safety review, while limiting its most sensitive cyber functions to verified users.
Z .ai计划在两周安全审查后发布GLM-5.3的权重,同时将其最敏感的网络安全功能限制给经过验证的用户。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力