跳到主内容
@wquguru
精选86Ars Technica AI(RSS)模型发布/更新

AI水印可能削弱大模型对抗恶意提示的安全性

LLMs respond differently to harmful prompts when AI watermarking is used

原文
发到 X
推荐理由

直接关联Claude等主流模型的安全边界,提醒开发者和安全团队在部署合规水印前必须重新评估对抗攻击风险。

In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it.

为响应一项新的欧盟法律,AI平台正在实施用于对其生成内容进行水印标记的新方案。Anthropic最近披露,其未来的Claude模型将使用SynthID-Text,这是Google创建并以开源形式发布的一种方法。它使用一个密钥,微妙地改变模型在句子中选择下一个词的过程。例如,原本最可能的下一个词选择可能是“cloudy”(多云),但密钥可能会将其改为“overcast”(阴天)。任何知道该密钥的人都可以判断内容是否由使用该密钥的平台生成。

New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information. Instructions that normally wouldn’t be followed will, in some cases, be performed once the watermarking is deployed. The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place.

最新研究表明,SynthID-Text不仅能改变词汇选择,还能改变模型调用的工具,以及模型遵循或无视其训练所应遵守的安全护栏的可能性。在面对对抗性提示时,这种威胁可能变得更大,攻击者试图诱导模型执行有害操作,例如泄露密码或其他敏感信息。在某些情况下,一旦部署了水印技术,那些通常不会被遵循的指令也会被执行。这一发现强调了开发人员需要彻底测试其在启用水印后大型语言模型(LLM)和智能体的行为表现。

Changing safety behavior

改变安全行为

“As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they’re powering an agent,” Andrea Siposova, an AI security researcher at Lasso Security, told Ars. “Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it’s going to show up somewhere.”

Lasso Security的AI安全研究员Andrea Siposova告诉Ars:“与没有水印的同款模型相比,这肯定会改变它们的行为,特别是当我们将其置于对抗性条件下,或让这些模型在驱动智能体时调用工具。”她补充道:“水印的设计初衷是使读者无法察觉,但我们知道,当我们要改变模型生成的任何内容时,都会产生一些权衡,并且这种影响会在某处显现出来。”

Read full article

阅读全文

Comments

评论

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件