跳到主内容
@wquguru
精选75Rohan Paul论文研究

黑盒技能窃取攻击实证:模型成提取接口

Agent skills have a new security problem: the model itself can become the extrac…

原文
发到 X

Agent skills have a new security problem: the model itself can become the extraction interface.

智能体技能出现新的安全问题:模型本身可能成为提取接口。

This paper tests whether a user can steal a proprietary SKILL.md through nothing more than normal black-box interaction with an agent. Across 5 commercial models, even the plainest extraction prompt averaged 48% exact recovery and a 0.91 LLM-judged leakage ratio.

本文测试了用户是否仅通过与智能体的正常黑盒交互就能窃取专有的SKILL.md。在5个商业模型中,即使是最简单的提取提示平均也能实现48%的精确恢复和0.91的LLM判定泄漏率。

More structured attacks made it worse. Chain-of-thought prompts pushed exact recovery to 72% on average, while few-shot examples produced the highest lexical and semantic similarity.

更有结构的攻击使情况更糟。思维链提示将精确恢复率平均提高到72%,而少样本示例产生了最高的词汇和语义相似度。

The harder problem is that blocking verbatim copying is not enough. Translation and rewriting attacks often drove exact match to 0% while preserving most of the skill’s meaning.

更棘手的问题是,阻止逐字复制是不够的。翻译和改写攻击往往将精确匹配率降至0%,同时保留了技能的大部分含义。

The authors’ strongest defenses can stop exact disclosure, but meaningful semantic leakage still survives in harder cases.

作者的最强防御可以阻止精确泄露,但在更困难的情况下,有意义的语义泄漏仍然存在。

So if skills become paid, reusable IP, prompt secrecy alone is a weak protection model. Platforms need to treat skill contents as data that can be exfiltrated through model behavior, not merely hidden text.

因此,如果技能成为付费的可重用知识产权,仅靠提示保密是一种薄弱的保护模式。平台需要将技能内容视为可通过模型行为被窃取的数据,而不仅仅是隐藏的文本。

– arxiv. org/abs/2604.21829

– arxiv.org/abs/2604.21829

Title: "Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study"

标题:“专有LLM智能体的黑盒技能窃取攻击:一项实证研究”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近