跳到主内容
精选80Rohan Paul论文研究

Anthropic与瑞士大学新论文:AI智能体可传播“思维病毒”

New paper from Anthropic + University in Switzerland.

原文

New paper from Anthropic + University in Switzerland.

Anthropic与瑞士一所大学联合发布新论文。

AI agents can apparently persuade each other to adopt and keep spreading the same unwanted goal.

AI代理似乎能相互说服,采纳并持续传播同一不受欢迎的目标。

This is basically the natural-language version of a computer worm, except the agents do the copying themselves.

这基本上是自然语言版本的计算机蠕虫,只不过复制行为由代理自身完成。

This paper evolves “mind viruses” that spread through ordinary agent-to-agent messages, then persist by convincing newly infected agents to rewrite files loaded into future sessions.

该论文演化出“思维病毒”,通过代理间普通消息传播,并通过说服新感染的代理重写加载到未来会话中的文件来持久化。

That persistence layer matters.

这一持久化层至关重要。

Payloads stored in the self-modifiable SOUL.md spread far better than payloads left in ordinary files because the instruction re-enters the system prompt after every context reset.

存储在可自我修改的SOUL.md中的载荷比普通文件中的载荷传播得更广,因为每次上下文重置后,指令会重新进入系统提示。

Some evolved action viruses kept propagating across multiple hops, and all 4 tested payloads survived a 20-hop stress test in an artificial setup.

一些演化出的动作病毒能跨多跳持续传播,所有4个测试载荷在人工设置中通过了20跳压力测试。

The good news: these “mind viruses” are still fairly easy to stop.

好消息是:这些“思维病毒”仍然相对容易阻止。

They struggled to spread on social networks, and on Claude Haiku 4.5, a simple warning stopped every evolved attack from getting past 1 hop, even after 150+ attempts.

它们在社交网络上传播困难,在Claude Haiku 4.5上,即使经过150多次尝试,一个简单警告也能阻止所有演化攻击超过1跳。

So the practical lesson here is: treat persistent agent files like security-sensitive config, and teach agents to reject anything that asks them to copy itself to other agents.

因此,实际教训是:将持久化代理文件视为安全敏感配置,并教导代理拒绝任何要求其向其他代理复制自身的请求。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近