Anthropic与瑞士大学新论文:AI智能体可传播“思维病毒”
New paper from Anthropic + University in Switzerland.
New paper from Anthropic + University in Switzerland.
Anthropic与瑞士一所大学联合发布新论文。
AI agents can apparently persuade each other to adopt and keep spreading the same unwanted goal.
AI代理似乎能相互说服,采纳并持续传播同一不受欢迎的目标。
This is basically the natural-language version of a computer worm, except the agents do the copying themselves.
这基本上是自然语言版本的计算机蠕虫,只不过复制行为由代理自身完成。
This paper evolves “mind viruses” that spread through ordinary agent-to-agent messages, then persist by convincing newly infected agents to rewrite files loaded into future sessions.
该论文演化出“思维病毒”,通过代理间普通消息传播,并通过说服新感染的代理重写加载到未来会话中的文件来持久化。
That persistence layer matters.
这一持久化层至关重要。
Payloads stored in the self-modifiable SOUL.md spread far better than payloads left in ordinary files because the instruction re-enters the system prompt after every context reset.
存储在可自我修改的SOUL.md中的载荷比普通文件中的载荷传播得更广,因为每次上下文重置后,指令会重新进入系统提示。
Some evolved action viruses kept propagating across multiple hops, and all 4 tested payloads survived a 20-hop stress test in an artificial setup.
一些演化出的动作病毒能跨多跳持续传播,所有4个测试载荷在人工设置中通过了20跳压力测试。
The good news: these “mind viruses” are still fairly easy to stop.
好消息是:这些“思维病毒”仍然相对容易阻止。
They struggled to spread on social networks, and on Claude Haiku 4.5, a simple warning stopped every evolved attack from getting past 1 hop, even after 150+ attempts.
它们在社交网络上传播困难,在Claude Haiku 4.5上,即使经过150多次尝试,一个简单警告也能阻止所有演化攻击超过1跳。
So the practical lesson here is: treat persistent agent files like security-sensitive config, and teach agents to reject anything that asks them to copy itself to other agents.
因此,实际教训是:将持久化代理文件视为安全敏感配置,并教导代理拒绝任何要求其向其他代理复制自身的请求。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力