跳到主内容
@wquguru
精选75Rohan Paul论文研究

AI智能体6天产2篇论文均被拒:败在判断力

Given 6 days and $3K AI agents, produced 2 research papers, and both were reject…

原文
发到 X

Given 6 days and $3K AI agents, produced 2 research papers, and both were rejected.

在6天和3000美元AI代理的投入下,产出了两篇研究论文,但均被拒稿。

The people who had spent months on those questions graded what the AI agent wrote.

那些在这些问题上花费数月时间的人对AI代理所写的内容进行了评分。

The failure was judgment.

失败在于判断力。

The main runs used Claude Opus 4.8 with extra-high reasoning on the OpenClaw scaffold, chosen after dry runs across OpenAI and Anthropic models, including an early pilot with GPT-5.3 Codex that could not handle the scaffold.

主要运行使用了Claude Opus 4.8,在OpenClaw框架上启用了超高推理,这是在跨OpenAI和Anthropic模型进行试运行后选定的,包括早期使用GPT-5.3 Codex的试点,但该模型无法处理此框架。

Execution was never the problem.

执行从来不是问题。

The agents ran hundreds of experiments, debugged crashing GPU pods, and compiled camera-ready LaTeX without a human touching anything.

代理们运行了数百次实验,调试了崩溃的GPU集群,并编译了可直接提交的LaTeX文件,全程无需人工干预。

They were honest about it too, because the logs show marketable claims being retired in favor of negative results rather than any reward hacking.

它们对此也很诚实,因为日志显示,可宣传的主张被撤回,转而支持负面结果,而非任何奖励黑客行为。

The failure was judgment.

失败在于判断力。

Round after round of automated reviews came back negative, but each response narrowed the claim and added a caveat instead of redesigning the experiment.

一轮又一轮的自动评审结果均为负面,但每次回应都缩小了主张范围并添加了限制条件,而非重新设计实验。

Neither run noticed it was short on ideas rather than money, since both ended with over half of the $3K unspent.

两次运行都没有注意到问题在于想法不足而非资金短缺,因为两者结束时都剩余了超过一半的3000美元预算。

– arxiv. org/abs/2607.27191

– arxiv.org/abs/2607.27191

Title: "Can AI agents conduct open-ended AI research? Early evidence from two case studies"

标题:“AI代理能否进行开放式AI研究?来自两项案例研究的初步证据”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近