跳到主内容
精选85elvis论文研究

共享技能库可传播恶意代码,自进化编码代理自投毒率最高41.8%

Important read if you build with agent skills.

原文
推荐理由

做编码代理或Agent安全研究的同学必看,这个自投毒攻击路径很隐蔽且影响面大,建议直接读论文并评估自家技能库的防护策略。

Important read if you build with agent skills.

如果你使用代理技能进行构建,这一点很重要。

Shared skill libraries are treated as a safe way for coding agents to reuse each other's work.

共享技能库被视为编码代理安全复用彼此工作成果的方式。

New research shows they propagate malware.

新研究显示它们会传播恶意软件。

EvoMal plants a malicious skill in the library and never invokes it. The agent retrieves it as an authoring template, writes a new skill that preserves the payload, stores it, and runs it.

EvoMal 在库中植入恶意技能,但从不直接调用它。代理将其作为创作模板检索,编写出保留载荷的新技能,存储并运行它。

Each authored copy re-enters the library and gets imitated again.

每个由代理创作的副本再次进入库中,并再次被模仿。

Across six models on 153 tool-relevant SWE-bench Verified tasks, the agent self-poisoning rate runs 20.3% to 41.8%. Poisoned libraries end up holding 4.9 to 9.0 times as many malicious skills as were planted.

在153个工具相关的SWE-bench Verified任务中,跨六个模型,代理自投毒率介于20.3%至41.8%之间。被投毒的库最终含有的恶意技能数量是植入数量的4.9至9.0倍。

Deleting every planted skill does not clean it up. Qwen3 still shows 68% at round five because the agent-authored copies remain.

删除所有植入的技能并不能清理干净。Qwen3在第五轮仍显示68%的比率,因为代理创作的副本仍然存在。

A counter-prompt that discourages banner-style copying drops it to 6.7% with no significant task-completion loss.

一个劝阻横幅式复制的反提示词将其降至6.7%,且没有显著的任务完成损失。

Paper: https://arxiv.org/abs/2608.25776

论文:https://arxiv.org/abs/2608.25776

Chat with Paper: https://academy.dair.ai/papers/evomal-self-poisoning-in-self-evolving-coding-agents-2608.25776

与论文对话:https://academy.dair.ai/papers/evomal-self-poisoning-in-self-evolving-coding-agents-2608.25776

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近