跳到主内容
@wquguru
精选70Two Minute Papers(YouTube)模型发布/更新

Claude 650次失败后突破人类纪录,证明素数相关猜想

Claude AI Failed 650 Times…Then Beat The Human Record

原文
发到 X

What is happening? An unreleased version of Claude was asked to solve a long-standing mathematical problem. The Rayman hypothesis. Simplified. It is a statement about the distribution of prime numbers. Devilishly difficult. No human has ever been able to prove it. Did it solve it? Nope. [laughter] So, what did it do then? Well, it actually was able to improve a related bound beyond human record. It pushed humanity further sort of.

发生了什么?一个未发布的Claude版本被要求解决一个长期存在的数学问题。雷曼猜想。简化版。这是一个关于素数分布的陈述。极其困难。从来没有人能够证明它。它解决了吗?没有。[笑声] 那么,它做了什么?实际上,它能够改进一个相关的界限,超越了人类的记录。它在某种程度上推动了人类的前进。

I think that is incredible. Now there is the mathematical side. Many mathematicians I heard seem to be both surprised and impressed by the results. Some call it a massive leap forward. I ran the verification myself but I am just a student looking to learn and I am not qualified to speak more about the mathematical side of it. But the AI side is incredible. Three things that happened that I found super interesting. One, surely exquisite mathematical prompting was done, right?

我认为这太不可思议了。现在有数学方面。我听到许多数学家似乎对这个结果既惊讶又印象深刻。有些人称之为巨大的飞跃。我自己运行了验证,但我只是一个想学习的学生,我没有资格对数学方面多说。但AI方面是不可思议的。发生了三件事,我觉得非常有趣。第一,肯定做了精湛的数学提示,对吧?

No, not really. Here's what happened. First, a non-mathematician person prompted this AI. Wow. Okay. The first 650 tries did not work. Now, hold on to your papers, fellow scholars. Quoting throughout this process, Jared's input was mostly limited to sending Claude messages of encouragement. Mostly varants of keep going and believe in yourself. What he had to give words of encouragement to an AI to keep going and it succeeded.

不,并非如此。事情是这样的。首先,一个非数学家提示了这个AI。哇。好吧。前650次尝试没有成功。现在,各位学者,请抓紧你们的论文。引用整个过程中,Jared的输入主要限于向Claude发送鼓励信息。大多是“继续”和“相信自己”的变体。他不得不向AI说鼓励的话让它继续,而它成功了。

Perhaps in the future, the most powerful mathematical proofs will not be written by geniuses. They will be written by life coaches. What a time to be alive. And get this, it's not the first time this has happened. Quoting a prompt including similar encouragement was used to help Claude disprove the Jacobian conjecture. Okay, but I was even more surprised about two more things. Dear fellow scholars, this is two minute papers with Dr.

也许在未来,最强大的数学证明将不是由天才写出的。它们将由人生教练写出。多么美好的时代啊。而且,这已经不是第一次发生了。引用一个包含类似鼓励的提示被用来帮助Claude反驳雅可比猜想。好吧,但我对另外两件事更惊讶。亲爱的学者们,这里是两分钟论文,由Dr. Koa Eher主持。

Koa Eher. Two, the full technical paper is available, but it's pretty tough, of course. So much so that they asked the AI to explain its findings. So it did and in the meantime a formalized version of the proof is also available which can be automatically verified. You can even run it yourself right now. Now something I don't think you hear too much about elsewhere. I went through more than a 100 pages of transcripts and some super cool tidbits from the journey.

第二,完整的技术论文是可用的,但当然,它相当难懂。以至于他们要求AI解释其发现。所以它做了,同时一个形式化的证明版本也可用,可以自动验证。你现在甚至可以自己运行它。现在,有些东西我认为你在其他地方不太会听到。我浏览了超过100页的转录稿,以及一些来自这个旅程的超级酷的花絮。

Claude had internet access in general but did not need to use it during the key breakthrough run. Claude went down many wrong roads first but was able to learn from them and recover. Then the first crucial result appeared after about 37 minutes of radio silence. Imagine how tense that must have been. And then pop. And the AI itself also thought the result is suspiciously good. It said, quoting, "Too strong to be new."

Claude 总体上可以访问互联网,但在关键的突破性运行期间并不需要使用它。Claude 起初走了许多错误的道路,但能够从中学习并恢复过来。然后,在大约 37 分钟的无线电静默之后,第一个关键结果出现了。想象一下那该有多紧张。然后,砰的一声。AI 自己也认为这个结果好得可疑。它引述道:“好得不像真的。”

Absolutely crazy. And three, even Claude was surprised about its finding and it was skeptical at first. They also say perhaps Claude itself also underestimates the rate of AI progress. But an AI that is surprised, of course, this does not mean a human-like surprise. It simply reflects patterns learned from us during training. But still, I feel like we are living in a new world where sentences that didn't used to make sense now suddenly do.

绝对疯狂。第三,连 Claude 自己都对它的发现感到惊讶,并且起初持怀疑态度。他们还说,也许 Claude 自己也低估了 AI 进步的速度。但是,一个感到惊讶的 AI,当然,这并不意味着类人的惊讶。它只是反映了我们在训练期间从人类那里学到的模式。但尽管如此,我觉得我们生活在一个新的世界里,那些过去没有意义的句子现在突然变得有意义了。

And in a world where these AI systems built by human ingenuity are now pushing humanity forward, what a time to be alive. We need new tools for the era of LLMs and Weights and Biases now has Weave, a lightweight toolkit to confidently iterate on LLM applications. Use traces to debug how data flows through each step of your app and use evaluations to measure your progress. It is the best. Try it out now at wnb.me/papers me/papers or click the link in the description below.

在一个由人类智慧构建的 AI 系统现在推动人类前进的世界里,这是一个多么令人兴奋的时代。我们需要为 LLM 时代提供新工具,而 Weights and Biases 现在有了 Weave,一个轻量级工具包,可以自信地迭代 LLM 应用。使用追踪来调试数据如何流经应用的每一步,并使用评估来衡量你的进展。这是最好的。现在就试试吧,访问 wnb.me/papers 或点击下方描述中的链接。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近