MerchantBench:评估LLM Agent长期连贯性的电商基准
Most agent benchmarks end after one task, but running a store doesn't, and that'…
Agent评测常忽略长期稳定性,这篇论文提出的MerchantBench及活跃度监控指标对构建持久化Agent极具参考价值,值得收藏。
Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart.
大多数智能体基准测试在完成一个任务后就结束了,但经营商店并非如此,而这正是这些智能体拉开差距的地方。
MerchantBench hands an agent a small online store and lets it run for a simulated year, sourcing products, setting prices, managing cash. It also scores something most evals skip: whether the agent is still acting at all.
MerchantBench 向智能体提供一个小型在线商店,并让其运行模拟的一年时间,包括采购商品、设定价格和管理现金流。它还评估了大多数评测所忽略的一项指标:智能体是否仍在持续运作。
The best agent finished a simulated year of shopkeeping with about a quarter of what humans earned, mostly by going quiet, so track how often yours still acts.
表现最佳的智能体在模拟一年的店铺经营后,收入约为人类收入的四分之一,这主要是因为它后来变得“沉默”了,因此请跟踪你的智能体保持活跃的频率。
So log your agent's actions per window and watch that curve. A decent final score can hide an agent that stopped working months ago.
因此,记录每个时间窗口内智能体的操作次数并观察该曲线。一个看似不错的最终得分可能会掩盖智能体早在几个月前就已停止工作的真相。
– arxiv. org/abs/2607.28956
– arxiv.org/abs/2607.28956
Title: "MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations"
标题:《MerchantBench:评估电商运营中 LLM 智能体的长期一致性》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力