跳到主内容
精选85Rohan Paul论文研究

MDA框架:LLM提假设,贝叶斯定证据,实验效率大增

LLMs can suggest scientific mechanisms, but this paper finds that letting the ag…

原文
推荐理由

做科学发现和Agent研究的同学必看,MDA把LLM从决策者降为假设生成器,用贝叶斯推理和信息价值选实验,8次实验就达到41次的效果,赶紧读原文复现一下。

LLMs can suggest scientific mechanisms, but this paper finds that letting the agent choose experiments and fit the mechanism is far less data-efficient.

大语言模型可以提出科学机制,但本文发现,让智能体自行选择实验并拟合机制,其数据效率远低于预期。

So scientific agents may work better when the LLM proposes hypotheses but does not decide what the evidence means.

因此,当大语言模型提出假设但不决定证据含义时,科学智能体可能表现更好。

MDA turns the LLM into the hypothesis generator. Bayesian inference scores the candidate mechanisms, and value-of-information chooses the next experiment where those mechanisms disagree most.

MDA将大语言模型转化为假设生成器。贝叶斯推断对候选机制进行评分,信息价值则选择下一个实验,在这些机制分歧最大的地方进行。

That changes the experiment budget dramatically.

这极大地改变了实验预算。

On FORCEBENCH, MDA reaches roughly the accuracy of an unthrottled Opus 4.7 agent using 8 experiments instead of about 41, while reaching a 93% numeric pass rate versus 31% for the budget-matched Opus 4.7 LLM agent.

在FORCEBENCH上,MDA使用8次实验达到了与未限流的Opus 4.7智能体大致相当的准确度,而后者需要约41次实验;同时,MDA的数值通过率达到93%,而预算匹配的Opus 4.7大语言模型智能体仅为31%。

The mechanism is easy to see in the examples: for Yukawa forces, it probes long range because the competing laws look identical nearby; for Coulomb, it changes source charge because moving the probe alone cannot separate the true law from a charge-blind fit.

机制在示例中显而易见:对于Yukawa力,它探测长程效应,因为竞争定律在近距离处看起来相同;对于Coulomb力,它改变源电荷,因为仅移动探针无法将真实定律与无视电荷的拟合区分开。

When predictions still fail, MDA asks the LLM for new mechanisms and repeats the loop.

当预测仍然失败时,MDA会要求大语言模型提出新机制,并重复循环。

Let LLMs propose scientific ideas, but let explicit uncertainty and designed experiments decide what survives.

让大语言模型提出科学想法,但让明确的的不确定性和设计的实验来决定哪些想法能存活。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近