OpenAI用Astra模型攻克十道十年未解数学难题
Ten advances in mathematics and theoretical computer science
做AI研究或关注模型能力的同学必看,OpenAI用下一代模型Astra在数学推理上取得突破,成本极低,值得关注其技术细节和Lean 4形式化验证。
Ten advances in mathematics and theoretical computer science
A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings."
Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one.
(No news on how many problems they spent $2,000 on without reaching a solution though.)
The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces.
That's a decent level of transparency, but I want to see the prompts they used!
A lot of mathematicians online are experiencing a collective burst of Deep Blue.
OpenAI's results reminds me of what Terence Tao described as "big mathematics" in IEEE Spectrum in June:
Unlike some of his peers, Tao is neither dismissive of AI nor fearful. Instead, he sees it as the catalyst for a fundamental shift in the discipline—a transition toward what he calls “big mathematics.” He envisions a future of large-scale, decentralized collaborations between humans and machines, where complex mathematical tasks can be diced and sliced, with humans claiming the creative parts and AI doing the lion’s share of the technical grunt work.
Via Hacker News
Tags: mathematics, ai, openai, generative-ai, llms, deep-blue
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力