自主智能体跨迭代记忆使安全系统失效
If an autonomous agent remembers across iterations, its safety system cannot aff…
If an autonomous agent remembers across iterations, its safety system cannot afford to forget between them..
如果自主代理在多次迭代中保持记忆,其安全系统就不能在迭代之间遗忘。
long-horizon agent safety cannot be reduced to repeatedly running a safe short-horizon trajectory.
长期代理的安全不能简化为反复运行安全的短期轨迹。
An attacker can split malicious evidence across several individually benign-looking steps, so a trajectory-only monitor never sees enough context to distinguish the attack from normal work.
攻击者可以将恶意证据分散在多个看似无害的步骤中,因此仅基于轨迹的监控器永远无法获得足够的上下文来区分攻击与正常操作。
– arxiv. org/abs/2608.27141
– arxiv. org/abs/2608.27141
Title: "Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents"
标题:“安全不构成:自主LLM代理的非衰减循环状态”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力