GLM-5.3仅靠后训练大幅提升能力,权重两周后发布
GLM-5.3 shows how much capability may still be hiding inside today’s largest bas…
GLM-5.3 shows how much capability may still be hiding inside today’s largest base models and how relevant post-training really is.
It uses the same base model (!) as GLM-5.2. Zai says the entire (!) improvement came from scaling post-training: more executable environments, longer tasks, stronger verifiers and more reinforcement learning.
Remember: Pre-training gives a model knowledge and raw problem-solving capacity. Post-training teaches it how to use that capacity: plan, call tools, test solutions, recover from failure and complete work over long horizons.
In cyber evaluations, GLM-5.3 moved from 24.4% to 54.4% on ExploitBench and completed 105 ExploitGym tasks in two hours, up from 29 for GLM-5.2. .
Its weights are scheduled for release in two weeks. However, numerous other open weight models will be released in the coming weeks:
-DeepSeek v4 Pro -Qwen3.8 27b -LTX 2.5 -Nemotron-Lighting -DeepSeek harness (just released, but harness isntead of a model) -Muse-Glimmer-30B (just released)
to name a few.
The US has meanwhile created classified cyber benchmarks and a voluntary pre-release process for "covered frontier models." What this release shows me, first and foremost, is that open models are continuing to move closer and closer to Frontier. And therefore, I believe that the US government will now further expand the regulatory framework to include open models.
That's why I'm even more excited for the ChatGPT "Astra" release. Because this model is *also* receiving a new (and more extensive) pre-training component, and we're currently seeing how much additional capability is enabled through post-training.
That's why this release is so significant; it demonstrates just how many areas for improvement are possible.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力