跳到主内容
精选88Two Minute Papers(YouTube)模型发布/更新

智谱发布 GLM-5.3 系列:开源权重,性能逼近闭源旗舰

GLM 5.3: Powerful AI Is Becoming Almost Free

原文
推荐理由

GLM-5.3 系列以开源权重实现了逼近顶级闭源模型的性能,且架构创新显著降低了推理成本,对开发者极具参考价值。

I am running this AI at home and it can do light simulations computing a beautiful image like this. [laughter] Look at that yummy costics too. Oh my. It can write up a fun strategy game. It can model a beautiful full 3D scene in free and open-source software like Blender. It is incredibly good and anyone can have all this at the grand total of let me see nothing. Yes, it's free and open weights. This is called GLM 5.3 Flash plus its big brother GLM 5.3.

我在家运行这个 AI,它就能进行轻量级的模拟计算,生成像这样精美的图像。[笑声] 看看那诱人的成本效益。天哪。它能编写有趣的策略游戏。它能在 Blender 等免费开源软件中建模出美丽的完整 3D 场景。它非常出色,而且任何人都可以以总计——让我想想——零成本的代价拥有这一切。是的,它是免费的且权重开源。这被称为 GLM 5.3 Flash,以及它的哥哥 GLM 5.3。

Flash was first stealthily released under a different name. And I mean [laughter] what? Look at that. It quickly overtook even deepseek in usage. I understand novelty and all, but this is crazy. And now hold on to your papers, fellow scholars, because if you let it think for a while, both the bigger and smaller models can get up close to fable level on some benchmarks. In my opinion, this does not feel like fable level in general, but in some experiments, it is not that far away.

Flash 最初是悄悄以另一个名字发布的。我是说 [笑声] 什么?看那个。它在用量上迅速超越了甚至包括 deepseek 在内的其他模型。我理解新奇感之类的因素,但这简直疯狂。现在请扶好你们的论文,各位学者同行,因为如果让它思考一段时间,较大和较小的模型在某些基准测试上都能接近 fable 级别。在我看来,总体而言这并不感觉像 fable 级别,但在某些实验中,差距并不大。

I can't believe that I am saying this, but I think within this year, in a few months, Fable might be surpassed by free AI systems. What a time to be alive. Okay, so how is this black magic possible? How did they do it? Dear fellow scholars, this is 2 minute papers with Dr. Koa. Well, GLM 3.5 flash is about cramming more intelligence into less compute. 320 billion parameters about 95% of which is not activated per token.

我不敢相信我会这么说,但我认为在今年之内,几个月内,fable 可能会被免费 AI 系统超越。生逢此时真是幸事。好了,那么这种黑魔法是如何成为可能的?他们是怎么做到的?亲爱的学者同行们,这是由 Koa 博士带来的两分钟论文。嗯,GLM 3.5 flash 大约是将更多的智能塞入更少的计算资源中。它有 3200 亿个参数,其中约 95% 在每个 token 时未被激活。

The number of layers was 92 before. Now it's roughly cut in half. And it has some new tricks up the sleeve. Instead of comparing every token to every other one, it can summarize nearby context into a small package. They call this linear attention and [clears throat] it is dramatically cheaper than sparse attention which they also use. Also ever experience talking to an AI in a session for a long while and as you talk longer and longer it gets worse and worse.

之前的层数是 92 层。现在大约减半了。它还藏着一些新招数。它不是将每个 token 与其他所有 token 进行比较,而是可以将附近的上下文总结为一个小的包。他们称之为线性注意力(linear attention),并且 [清嗓子] 这比他们也使用的稀疏注意力(sparse attention)要便宜得多。大家是否也有过在会话中长时间与 AI 对话的经历,随着对话越来越长,表现变得越来越差?

Yes, if you have a bunch of documents, a huge codebase or a long discussion, searching through that context gets really expensive. And this can compress an index of the stored context before searching it. So the model can look back far further and use less memory and compute. They call this index pool. And if you add this all together, you have a system that is built from the ground up to be fast, smart, and inexpensive.

是的,如果你有一堆文档、一个巨大的代码库或一场漫长的讨论,搜索这些上下文会变得非常昂贵。而这可以在搜索之前压缩已存储上下文的索引。因此,模型可以回溯得更远,并使用更少的内存和计算资源。他们称之为索引池(index pool)。如果将这些全部结合起来,你就拥有一个从底层构建旨在快速、智能且廉价的系统。

Now, it still needs beefy hardware in the order of thousands of dollars. But what I love about you brilliant fellow scholars is that you are always tirelessly working on making it run on more modest hardware. This is once again an openweight system, a project that we should work on together and add to it. So thank you so much to all of you who contribute by improving it, anyone who runs experiments with it to share with us or just by talking about it and spreading the word.

现在,它仍然需要价值数千美元的强劲硬件。但我最喜欢你们这些才华横溢的学者的一点是,你们总是孜孜不倦地致力于让它能在更普通的硬件上运行。这再次是一个开源系统,一个我们应该共同努力并加以完善的项目。因此,非常感谢所有通过改进它、用它进行实验并与我们分享,或者仅仅通过谈论它来传播消息的贡献者。

We have some of the most amazing audience on this channel and it is an honor to make these videos for you. This is paper video number 1,072 and I'm having more fun with you than ever. This is why I said no to millions of dollars from private equity firms that wanted to buy our channel. Every single other one they approached I know and heard about was sold, but we didn't sell. Although [laughter] I would have a great deal better hardware, that's for sure.

我们拥有这个频道最精彩的观众,为你们制作这些视频是一种荣幸。这是第 1,072 期论文视频,我比以往任何时候都更享受与你们的互动。这就是为什么我拒绝了私募股权公司想要购买我们频道的数百万美元报价的原因。我知道并听说过的每一个其他被接触的频道都被出售了,但我们没有卖。虽然[笑声]我会拥有好得多的硬件,那是肯定的。

Yes, about that. The main limitation for me here is of course hardware. You need lots of firepower to run this. It's not in the billions anymore, which is fantastic, but it's still not cheap. I also don't have all the hardware to run the big one at home. Lambda helps though. So many of us will use smaller compressed quantized versions and these can start looping like crazy like it did for me here. So do not expect perfection but expect to have a damn good time tinkering and I am incredibly grateful for this.

是的,关于这一点。对我来说这里的主要限制当然是硬件。你需要大量的算力来运行它。它不再需要数十亿美元的资金了,这太棒了,但仍然不便宜。我也没有在家里运行大型模型的所有硬件。不过 Lambda 提供了帮助。所以我们会使用许多较小的压缩量化版本,这些版本可以像我在上面展示的那样疯狂地循环运行。所以不要期待完美,但要期待在折腾中获得极大的乐趣,我对此感激不尽。

Thank you so much. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to image or video. Easy peasy. Running a Deepseek chatbot or agent. Super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.

非常感谢。我经常在几分钟内使用 Lambda 复现 AI 研究论文。训练你自己的模型或微调现有模型也非常棒。运行推理或文本到图像或视频。轻而易举。运行 Deepseek 聊天机器人或智能体。超级快,超级可靠。Lambda 为你提供强大的 Nvidia GPU 来运行你自己的实验。我测试我所涵盖论文中的想法,片刻之后就能看到结果。非常喜欢。说真的,现在就试试 lambda.ai/papers。

Ei peepers.

再见啦。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近