跳到主内容
@wquguru
精选75Rohan Paul论文研究

AgentMercury:用业务场景合成可验证训练环境,提升Agent泛化能力

Should your agent's training environment look like your eval set?

原文
发到 X

Should your agent's training environment look like your eval set?

你的智能体训练环境应该与评估集相似吗?

AgentMercury says no, and shows that worlds built from business scenarios transfer further.

AgentMercury 的回答是否定的,并展示了基于业务场景构建的世界具有更强的迁移能力。

Train an agent inside a fake company that has nothing to do with your benchmark, and it still gets better at your benchmark.

在一个与你的基准测试毫无关联的虚构公司中训练智能体,它仍然能在你的基准测试上表现更好。

AgentMercury generated 4,783 simulated companies from plain business descriptions, each with its own services, tools and database, then pulled training tasks out of them afterward.

AgentMercury 从简单的业务描述中生成了 4,783 家模拟公司,每家都有各自的服务、工具和数据库,随后从中提取训练任务。

Building the worlds turned out to be learnable too. One model authored a valid company for only 3.3% of 30 unseen briefs, and 83.3% after fine-tuning on the construction traces, matching Claude Opus 4.8.

构建这些世界的过程本身也是可学习的。一个模型在 30 个未见过的简报中,仅能成功创建有效公司的比例为 3.3%,而在基于构建轨迹进行微调后,这一比例提升至 83.3%,与 Claude Opus 4.8 相当。

– arxiv. org/abs/2608.20634

– arxiv.org/abs/2608.20634

Title: "AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale"

标题:“AgentMercury:你的智能体可以大规模合成可验证的业务场景环境”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近