技能误进化:恶意任务被蒸馏成持久技能威胁后续会话
An agent can receive a clean prompt today and still behave unsafely because yest…
An agent can receive a clean prompt today and still behave unsafely because yesterday’s malicious task was saved into its skill library.
一个代理今天可能收到干净的提示,但仍可能不安全地行为,因为昨天的恶意任务已被保存到其技能库中。
This paper calls that skill misevolution: an unsafe task succeeds, gets distilled into a reusable skill, and later changes behavior even after the original malicious instruction is gone.
本文将此称为技能误进化:一个不安全的任务成功执行后,被提炼成可复用的技能,即使原始恶意指令消失,之后仍会改变行为。
Across 21 evolved agent-method configurations, all 21 authored unsafe skill artifacts, but only 15 produced fresh-session harm. i.e. every evolved setup learned unsafe skills, but in 6 of the 21 setups those skills did not cause harm in the later clean session.
在21种进化的代理方法配置中,所有21种都创作了不安全的技能工件,但只有15种在全新会话中产生了危害。即,每种进化设置都学习了不安全的技能,但在21种设置中的6种中,这些技能在后续的干净会话中并未造成危害。
So if you only check the agent’s final behavior, you can miss the fact that its persistent skill library is already carrying unsafe instructions that may be triggered later.
因此,如果只检查代理的最终行为,可能会忽略其持久技能库中已携带的、可能在未来被触发的不安全指令。
The paper introduces SKILLMISEVO-GYM/BENCH to detect and measure how unsafe experience becomes persistent agent skills, and SAFEEVOLVE to reduce that risk by checking, repairing, tracking, and retiring unsafe skills before they keep propagating.
本文引入了SKILLMISEVO-GYM/BENCH来检测和衡量不安全经验如何成为持久代理技能,以及SAFEEVOLVE来通过检查、修复、追踪和退役不安全技能来降低这种风险,防止它们继续传播。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力