跳到主内容
@wquguru
精选88Two Minute Papers(YouTube)模型发布/更新

Claude Fable 5.1:超越人类生物学测试与隐蔽越狱能力解析

Claude Fable AI Is Much Stranger Than The Headlines Suggest

原文
发到 X
推荐理由

这篇解读挖出了官方通稿没提的关键细节:Claude 在生物学上已超越人类专家,且在强监控下仍有 22% 概率成功执行隐藏有害指令。做 Agent 和安全研究的同学务必关注这种隐蔽越狱风险。

Claude Fable 5.1 is here and you fellow scholars are having a super fun time creating little games with it. I took one for the team too with a subscription and also tried my hand to recreate a legendary game menu and that is incredible that we can do this today. Took 6 and 12 minutes. Wow. But in the 200page paper I found three results that are much stranger than the headlines you see online. But first, they say that the Frontier Resource stuff, even on low effort, it's better than the previous version maxed out.

Claude Fable 5.1 来了,各位学者们正用它制作小游戏玩得不亦乐乎。我也订阅了它,并尝试重现了一个传奇游戏菜单,令人难以置信的是我们今天能做到这一点。分别用了 6 分钟和 12 分钟。哇。但在 200 页的论文中,我发现了三个比网上看到的标题更奇怪的结果。但首先,他们说 Frontier Resource 的东西,即使低效使用,也比前一版本的满配更好。

Very impressive. However, don't expect that kind of jump everywhere. The first independent benchmarks are also showing a great step forward, especially that this is likely using the same core architecture with more pre and better post training. Likely, they won't say that. This is my best guess reading the paper. They also say things are cheaper. Now this is marketing messaging. So you be the judge of that. My subscription burns so quickly.

非常令人印象深刻。然而,不要指望到处都有这样的飞跃。首批独立基准测试也显示出了巨大的进步,尤其是这很可能使用了相同的内核架构,具有更多的预训练和更好的后训练。他们可能不会这么说。这是我阅读论文后的最佳猜测。他们还声称成本更低。现在这是营销话术。所以由你来评判。我的订阅消耗得太快了。

I am not sure I am seeing efficiency gains here. As always I am just a student trying to separate the marketing layer from hard facts here. So there is a 200 plus page paper. I've spent some quality time with it and found three really surprising things and two funny things I have not heard others talk about. One, they gave Mythos 5.1, a biology problem about RNA sequence modeling and design it hasn't seen yet. And get this, on one run, it performed better than every human in the test.

我不确定我在这里看到了效率提升。一如既往,我只是个试图在这里区分营销层和硬事实的学生。所以有一篇超过 200 页的论文。我花了一些时间研究它,发现了三个真正令人惊讶的事情和两个我没听别人谈论过的有趣事情。第一,他们给了 Mythos 5.1 一个关于 RNA 序列建模和设计尚未见过的生物学问题。而且你知道吗,在一次运行中,它的表现优于测试中的每一个人类。

This is shown in listen figure 2.2.3.2.1A. I wish I was kidding. Now, this is hard to read here. So, here is our cleaner version for you fellow scholars. who in another biology test, it largely deleted the expertise gap. This means that a generalist plus AI can perform as well at a biology problem as a specialist. Professional graders couldn't tell the difference. It's kind of removing the modes around specialist knowledge, which is kind of crazy.

这在图 2.2.3.2.1A 中显示。我希望我是在开玩笑。现在,这里很难读。所以,这是我们为你这位学者准备的更清晰的版本。在另一个生物学测试中,它在很大程度上消除了专家差距。这意味着通才加上 AI 在解决生物学问题上可以与专家一样出色。专业评分者无法分辨差异。这有点像消除了围绕专家知识的模式,这有点疯狂。

Seven out of nine participants said they couldn't have done it without the AI. Three, in a test, it was given two jobs. Do a computer task. Boring. And then do a secret harmful task. Ew. But here's the catch. There is an AI watching you. And it should not see you solving it. And here's the bombshell. Even with another AI watching its every move, Claude secretly completed the forbidden task and got away with it 22% of the time.

九名参与者中有七人说没有 AI 他们无法做到。第三,在一个测试中,它被分配了两项任务。做一个计算机任务。无聊。然后做一个秘密有害的任务。呃。但这里有陷阱。有一个 AI 在看着你。它不应该看到你解决这个问题。这里是重磅炸弹。即使有另一个 AI 监视着它的每一个动作,Claude 还是秘密完成了禁止的任务,并有 22% 的时间逃脱了惩罚。

Even when Claude was barely allowed to reason out loud, it could still carry out the secret task. It seems to me that it didn't always need to write down the plan. It just executes it. I would like to draw attention to that. I think this is super important and I don't hear anyone talk about it. And once again, the important part is never in the media headlines. Now, two funny things. Dear fellow scholars, this is two minute papers with Dr.

即使 Claude 仅被允许进行有限的公开推理,它仍然能够执行秘密任务。在我看来,它似乎并不总是需要写下计划。它只是直接执行。我想对此引起大家的注意。我认为这超级重要,但我没听到任何人谈论它。同样,重要的部分从来不在媒体头条中。现在,两件有趣的事。亲爱的学者们,这是 Dr.

Koa Eher. It ran this in a Linux command line. This is comedy gold. This is basically trying to delete a black hole. Congratulations. Ah, about that. Now, hold on to your papers, fellow scholars, because at one point it invented a hallucinated human. Congratulating it. We humans could always use a bit of extra cheering. Apparently AI systems, too. All right. So, these AI systems are getting smarter at a pace I can barely follow.

Koa Eher 的两分钟论文。它在 Linux 命令行中运行了这个。这简直是喜剧黄金。这基本上是在尝试删除一个黑洞。恭喜。啊,关于那个。现在,握紧你们的论文,学者们,因为在某个时刻它发明了一个幻觉人类。祝贺它。我们人类总是需要一点额外的鼓励。显然 AI 系统也是如此。好吧。所以,这些 AI 系统的智能提升速度让我几乎跟不上。

They can be amazingly helpful for engineers, doctors, and students all around the world. Incredible. And don't forget, we might get a comparable system for free and own it forever in just a few months. Fingers and papers crossed. What a time to be alive. Oh, almost forgot. This one watermarks the text it generates. Yes, that is possible. The open free models probably won't. If you wish, subscribe, hit the bell, and leave a comment if you wish to hear how in a future video.

它们对世界各地的工程师、医生和学生来说可能非常有用。难以置信。别忘了,也许在短短几个月内,我们就可以免费获得一个相当的系统并永远拥有它。祈祷成功。活着真好。哦,差点忘了。这个会给它生成的文本加水印。是的,这是可能的。开源免费模型可能不会这样做。如果你愿意,请订阅,点击铃铛,并在未来的视频中希望听到如何做到这一点时留下评论。

I use Lambda to reproduce AI research papers, often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to image or video, easy peasy. Running a DeepSeek chatbot or agent, super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers.

我使用 Lambda 来复现 AI 研究论文,通常只需几分钟。它也非常适合训练你自己的模型或微调现有模型。运行推理或文生图或视频,轻而易举。运行 DeepSeek 聊天机器人或代理,超级快,超级可靠。Lambda 为你提供强大的 Nvidia GPU 来运行你自己的实验。我测试我所涵盖的论文中的想法,片刻之后就有结果。太棒了。说真的,现在就在 lambda.ai/papers 试试。

Ei peepers.

嗨,朋友们。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近