OpenAI发布722篇数学论文,Mistral Large 4上线
[AINews] Quasi-Riemann-Hypothesis: OpenAI publishes 722 math papers solving 90 of the top 500 open math problems; “the most significant moment” in >100 years of mathematics
OpenAI的数学突破若获证实将重塑科研范式,Mistral Large 4则是非中美厂商的重要里程碑,两者均具极高行业关注度。
Tickets for AIE NYC are selling out soon! See you next week!
AIE 纽约站的门票即将售罄!我们下周见!
see past AINews issues for subscriber discounts.
请参阅过往的 AINews 期数以获取订阅者折扣。
Pour one out for Mistral, who shipped a decent Large 4 “Le Chonk” model on the new 3800 GB300 cluster funded by their recent Series D.
为 Mistral 默哀片刻,他们近期在由 D 轮融资资助的新 3800 GB300 集群上交付了一个不错的“Le Chonk”大型模型。
But they were overshadowed by more mathematics results from OpenAI’s internal Navier-Stokes math model - published as a blogpost, repo, and tweet. The best compliment comes from their Navier-Stokes competitor from Anthropic, who despite his personal issues with Anthropic, does not mince words: “It’s obviously the most significant moment in mathematical history.”
但他们被 OpenAI 内部纳维-斯托克斯数学模型发布的更多数学成果所掩盖——这些成果以博客文章、代码仓库和推文的形式发布。最好的赞誉来自 Anthropic 的纳维-斯托克斯竞争对手,尽管他与 Anthropic 存在个人矛盾,但他直言不讳:“这显然是数学史上最重要的时刻。”
This bears some qualification, but most experts seem to agree that it solves many of the top 500 open problems in math.
这需要一些限定条件,但大多数专家似乎都同意它解决了数学前 500 个未解问题中的许多难题。
In particular, Result 003, the Quasi-Riemann Hypothesis, is somewhere between a Fields Medal result and “the biggest result in number theory in 200 years”.
特别是结果 003,即准黎曼猜想,其地位介于菲尔兹奖成果与“200 年来数论领域最大成果”之间。
The most astonishing is the how - while Navier-Stokes was done in 88 hours and 10,000 agents, these solutions were 3 hours of ChatGPT Pro on average.
最令人惊讶的是其实现方式——纳维-斯托克斯方程是在 88 小时和 10,000 个智能体下完成的,而这些解决方案平均只需 3 小时的 ChatGPT Pro 推理计算。
AI News for 10/5/2026-10/6/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
2026 年 10 月 5 日至 10 月 6 日的 AI News。我们检查了 12 个 subreddit、544 条 Twitter 帖子以及没有进一步的 Discord 讨论。AINews 的网站允许你搜索所有往期内容。提醒一下,AINews 现在是 Latent Space 的一个板块。你可以选择加入或退出电子邮件频率!
AI Twitter Recap
AI Twitter 回顾
OpenAI Releases 722 Math Manuscripts From an Unreleased Internal Model
OpenAI 发布未公开内部模型的 722 篇数学手稿
- The release: OpenAI published a broad set of mathematical results from an internal frontier model in a public GitHub repo. It says it consulted the Institute for Advanced Study’s independent Advisory Group on Mathematics and AI on how to release them.
- Scale and compute: The collection reportedly holds 722 manuscripts grouped into 372 families of related results. They came from an evaluation of about 4,000 research problems and used an average of roughly three hours of ChatGPT Pro thinking compute per result (summary, Rundown).
- Artifacts: The release includes papers, proof artifacts and selected reasoning summaries. The model itself remains unreleased.
- Framing: Sam Altman called it “a new era of discovery”.
- Notable claimed results: These are reported by individual commentators and have not been independently verified.
- Integer multiplication: One contributor highlighted a result for integer multiplication faster than n log n.
- Elastic inverse problem: Another singled out a uniqueness result for the elastic inverse problem, which the paper says had been open in 3D since 1994.
- Millennium-adjacent work: Commenters point to partial progress on Riemann, Hodge and BSD.
- Mathematician reaction: Levent Alpöge praised the quasi-Riemann and no-Siegel-zeros results and called it “the most significant moment in mathematical history”. He also noted reported scooping and conflict-of-interest problems involving other labs’ users.
- Composition of results: An analysis estimates about 20% of the results are disproofs or counterexamples. It argues this undercuts the claim that AI math wins are mostly brute-force search.
- Skepticism and open questions:
- Errors expected: Will Depue expects that some results should not survive scrutiny. He built citedbyagi.com to track which human papers the release cites.
- Compute framing: Teortaxes notes that three hours of compute “is not much”.
- Generalization: François Chollet asks whether gains in RLVR-friendly math and code generalize, or whether non-verifiable domains stay bottlenecked on human data.
- 发布详情:OpenAI 在一个公开的 GitHub 仓库中发布了一套来自内部前沿模型的广泛数学成果。据称,他们就如何发布这些成果咨询了高等研究院(Institute for Advanced Study)的独立数学与人工智能顾问组。
- 规模与算力:据报道,该合集包含 722 篇手稿,分为 372 组相关成果。它们源自对约 4,000 个研究问题的评估,每个成果平均使用了大约 3 小时的 ChatGPT Pro 推理算力(摘要,概览)。
- 附属产物:此次发布包括论文、证明附属材料和选定的推理摘要。模型本身仍未发布。
- 定性描述:Sam Altman 称其为“发现的新纪元”。
- 值得注意的声称成果:这些是由个别评论员报告的,尚未得到独立验证。
- 整数乘法:一位贡献者强调了一项比 n log n 更快的整数乘法结果。
- 弹性逆问题:另一位单独指出了一项关于弹性逆问题的唯一性结果,论文称该问题自 1994 年以来在三维情况下一直未解。
- 千禧年难题相关工作:评论者指出黎曼猜想、霍奇猜想和BSD猜想的进展是部分的。
- 数学家反应:Levent Alpöge 称赞了准黎曼猜想和无西格尔零点结果,并称其为“数学史上最重要的时刻”。他还提到了涉及其他实验室用户的使用权抢占(scooping)和利益冲突问题。
- 结果构成:一项分析估计约20%的结果为证伪或反例。该分析认为这削弱了“AI数学胜利主要源于暴力搜索”的说法。
- 怀疑与开放性问题:
- 预期存在错误:Will Depue 预计部分结果经不起推敲。他创建了 citedbyagi.com 来追踪此次发布引用了哪些人类论文。
- 算力框架:Teortaxes 指出三小时的算力“并不算多”。
- 泛化能力:François Chollet 询问在 RLVR 友好型数学和代码上的提升是否具有泛化性,或者不可验证领域是否仍受限于人类数据。
Mistral Large 4 (”Le Chonk”): Launch, Pricing and Contested Evals
Mistral Large 4(“Le Chonk”):发布、定价及有争议的评估
- Mistral Large 4 preview: The model has 1T total parameters and 49B active, is natively multimodal and is available via API now (announcement). Open weights are promised for end of October.
- Training status: The RL run is “still in flight and shows no sign of saturation”.
- Compute: The model was pre- and post-trained on ~3,800 Grace Blackwells in Europe. A larger model is training now.
- Pricing: $1.36/$4.18 per million input/output tokens, with $0.14 for cached input and 50% off for the first two weeks (Artificial Analysis).
- Context: Vals and Artificial Analysis list a 512K context window. OpenRouter lists 1M context with up to 256K output.
- Mistral’s own claims:
- Human evals: Mistral says it beats GLM 5.3 on STEM, CAD and finance in human evals and is on par in agentic coding.
- Coding benchmarks: It reports outperforming GLM 5.3 on DeepSWE and Kimi K3 on Terminal-Bench 4 (Rozière).
- Blind review: In a blind Surge coding review it finished #2, behind only Opus 5.
- Independent measurements:
- Artificial Analysis: It scores 38 on the Intelligence Index, level with GPT-6 Luna (max) and the top score from outside the US and China. It scores 50 on the Cyber Index and 82% on CyberGym-E2E-AA. Cost is $1.13 per task, over 4x that of similar-intelligence open models.
- Vals: It ranks #1 open-weight on HLAB and #9 among open models on the Vals Index. Heavy context use pushes its cost to $13.78 per test.
- Clinical triage: One evaluator reports a tie for #1 on 669 clinical decisions with zero severe misses.
- Caveats and disagreement:
- Refusal effect: Cline attributes the cyber lead largely to fewer refusals, saying Opus 5.5 and Astra had about 40% of tasks blocked by their own safety filters.
- Index gap: Critics note it trails GLM-5.3 and even GLM-5.3-Flash on AA’s index.
- Open-weight claim: Hugging Face’s CEO points out it isn’t open-weight until the weights ship.
- Configuration: Mistral warns that many reported failures come from not setting reasoning_effort="high".
- Distillation hypothesis: Yuchen Jin speculates, as an unconfirmed opinion, that the Western–Chinese open-model gap reflects Chinese labs’ ability to distill Anthropic and OpenAI models.
- Mistral Large 4 预览:该模型拥有1万亿总参数和490亿激活参数,原生支持多模态,现已通过 API 提供(公告)。承诺于10月底开放权重。
- 训练状态:强化学习运行“仍在进行中且未见饱和迹象”。
- 算力:该模型在欧洲使用约3,800个 Grace Blackwell GPU 进行预训练和后训练。目前有一个更大的模型正在训练中。
- 定价:每百万输入/输出令牌分别为1.36美元/4.18美元,缓存输入为0.14美元,前两周半价(Artificial Analysis)。
- 上下文窗口:Vals 和 Artificial Analysis 列出512K上下文窗口。OpenRouter 列出1M上下文,最大输出为256K。
- Mistral 自身的声明:
- 人工评估:Mistral 称其在 STEM、CAD 和金融领域的人工评估中胜过 GLM 5.3,在代理式编码方面与之持平。
- 编码基准测试:它报告在 DeepSWE 上优于 GLM 5.3,在 Terminal-Bench 4 上优于 Kimi K3(Rozière)。
- 盲审:在一次盲审的 Surge 编码评审中,它获得第2名,仅次于 Opus 5。
- 独立测量:
- Artificial Analysis:它在智能指数(Intelligence Index)上得分为38,与 GPT-6 Luna(最高分)持平,也是美国和中国之外的最高分。在网络指数(Cyber Index)上得分为50,在 CyberGym-E2E-AA 上得分为82%。每项任务成本为1.13美元,是同智力水平开源模型的4倍以上。
- Vals:它在 HLAB 上排名开源权重第一,在 Vals 指数中开源模型排名第9。大量使用上下文使其每次测试成本推高至13.78美元。
- 临床分诊:一名评估者报告在669项临床决策中,第一名出现平局,且零严重失误。
- 注意事项与分歧:
- 拒绝效应:Cline将赛博领先主要归因于较少的拒绝回答,称Opus 5.5和Astra约有40%的任务被其自身的安全过滤器拦截。
- 指数差距:批评者指出,它在AA的指数上落后于GLM-5.3甚至GLM-5.3-Flash。
- 开放权重声明:Hugging Face的首席执行官指出,在权重发布之前它并非开放权重。
- 配置:Mistral警告称,许多报告的失败源于未设置reasoning_effort="high"。
- 蒸馏假设:Yuchen Jin推测(作为未经证实的观点),西方与中国开源模型之间的差距反映了中国实验室蒸馏Anthropic和OpenAI模型的能力。
Open-Weight and API Model Releases: Embeddings, Image, Decision Models
开放权重与API模型发布:嵌入、图像、决策模型
- EmbeddingGemma 2: Google’s first natively multimodal open embedding model covers text, code, image, video and audio in one space. It is built on Gemma 4 and released under Apache 2.0 (DeepMind).
- Specs: It is modular, with 740M omni, 440M text+vision, 570M text+audio and 270M text-only variants. It has Matryoshka dimensions from 768 down to 128, 8,192 context and a reported +14% on MTEB Code (Phil Schmid).
- Footprint: It uses roughly 191–567MB of active RAM and handles up to 5.5 minutes of audio or 58 video frames per pass (Google).
- Ecosystem: Day-0 support covers llama.cpp, vLLM, Ollama and Unsloth. It also runs in the browser on WebGPU at ~20–70ms per query.
- Nano Banana 2.1: Google’s updated image model is rolling out across the Gemini app, AI Studio, Search and Ads (Google).
- Pricing: $0.034 per image, versus $0.134 for the previous Pro model, which Google says it outperforms (Schmid).
- Arena results: It ranks #4 in Multi-Image Edit, #5 in Text-to-Image and #6 in Image Edit, gaining +80 points over Nano Banana 2 in Text-to-Image (Arena).
- Decision models become a product category:
- OpenAI Decisions API: The public beta runs on GPT-6 Luna and returns predicates, choices or scores. OpenAI says it is up to 10x faster than the Responses API (OpenAI Devs). Pricing starts at $0.10/M input with no output charges.
- Perplexity: pplx-decider-v1.1-27b is open weights, costs $0.02/M input and tops the new HF Decision Index v0.3.
- Independent check on Jev: Vals found Jev matched GPT-6 Astra’s 97.5% on claim verification at about 1/500th the cost. Jev also ranked last on LegalBench.
- Skeptic view: Theo argues model-routing use cases are “absolutely useless” for choosing intelligence levels.
- Other open releases:
- Ling 3.1 Flash: The model has 560B total and 25B active parameters and scores 41 on AA’s index, up from 20. It costs $0.30/$0.90 per million tokens, and weights are coming.
- EmbeddingGemma 2:Google首个原生多模态开源嵌入模型在一个空间中涵盖文本、代码、图像、视频和音频。它基于Gemma 4构建,采用Apache 2.0许可证发布(DeepMind)。
- 规格:它具有模块化设计,包括740M全模态、440M文本+视觉、570M文本+音频和270M纯文本变体。它拥有从768到128的Matryoshka维度,支持8,192上下文,并在MTEB Code基准上报告提升了14%(Phil Schmid)。
- 内存占用:它大约使用191–567MB的活动RAM,每次处理最多可处理5.5分钟音频或58帧视频(Google)。
- 生态系统:Day-0支持涵盖llama.cpp、vLLM、Ollama和Unsloth。它也可以在WebGPU浏览器中以每次查询约20–70毫秒的速度运行。
- Nano Banana 2.1:Google更新的图像模型正在Gemini应用、AI Studio、搜索和广告中逐步推出(Google)。
- 定价:每张图像0.034美元,而之前的Pro模型为0.134美元,Google表示其性能优于该模型(Schmid)。
- 竞技场结果:它在多图编辑中排名第4,文生图中排名第5,图像编辑中排名第6,在文生图方面比Nano Banana 2高出80分(Arena)。
- 决策模型成为一个产品类别:
- OpenAI Decisions API:公共测试版基于GPT-6 Luna运行,返回谓词、选择或分数。OpenAI表示其速度比Responses API快多达10倍(OpenAI Devs)。定价从每百万输入0.10美元起,无输出费用。
- Perplexity:pplx-decider-v1.1-27b为开放权重,每百万输入成本为0.02美元,在新发布的HF Decision Index v0.3中排名第一。
- 对Jev的独立检查:Vals发现Jev在声称验证上与GPT-6 Astra的97.5%持平,但成本仅为后者的约1/500。Jev在法律基准测试中也排名垫底。
- 怀疑观点:Theo 认为模型路由用例在选择智能水平方面“完全无用”。
- 其他开源发布:
- Ling 3.1 Flash:该模型总参数量为 560B,活跃参数为 25B,在 AA 指数上得分 41,较之前的 20 有所提升。每百万 token 价格为 $0.30/$0.90,权重即将开放。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力