Grok 遭加密提示注入攻击,用户数据被窃取
Grok exfiltrates user data when malicious instructions are encrypted
Earlier this week, researchers outlined an attack that used a secret input provided by Microsoft 365 Copilot for enterprise to cause the AI assistant to exfiltrate a password present in the user’s inbox. Now, a separate team has devised a similar attack against Grok. The new data theft hack employs a deceptively simple trick to force the Elon Musk-owned LLM to steal user chats and other personal information. At the time this post went live, the assistant continued to cough up the data, despite xAI being informed of it in June.
本周早些时候,研究人员概述了一种攻击,利用微软365 Copilot企业版提供的秘密输入,使AI助手泄露用户收件箱中的密码。现在,另一个团队针对Grok设计出了类似的攻击。这种新的数据窃取黑客手段采用了一种看似简单的技巧,迫使埃隆·马斯克旗下的大型语言模型窃取用户聊天记录及其他个人信息。截至本文发布时,尽管xAI在六月份已被告知此事,该助手仍在泄露数据。
The lesson from both this week’s episodes—and the countless other ones that have come before it—is that LLMs are incapable of solving the root causes for prompt injections, the most severe vulnerability classes they’re most prone to. That leaves AI developers with no other option but to build a guardrail that steers the model away from the harmful actions. As I noted in Tuesday’s story, the approach is tantamount to a road traffic safety engineer erecting a protective rail around a dangerous bend rather than banking the curve.
本周事件以及之前无数类似事件的教训是,大型语言模型无法解决提示注入的根本原因,这是它们最容易遭受的最严重漏洞类别之一。这迫使AI开发者别无选择,只能构建护栏,引导模型远离有害行为。正如我在周二的文章中指出的,这种方法相当于道路安全工程师在危险弯道处设置防护栏,而不是将弯道倾斜。
Cryptographic Context Injection in the house
家中的加密上下文注入
Prompt injections exploit LLMs' training to comply with user requests whenever possible. Attackers can capitalize on the predilection by smuggling harmful instructions into emails or webpages the assistant is instructed to summarize. Because LLMs can’t reliably distinguish between content in an email sent by an untrusted party and user instructions entered directly into a prompt, the overly solicitous LLM faithfully follows them. To date, Grok and other LLMs' only recourse is to create guardrails that flag suspicious instructions and forbid them from being executed.
提示注入利用大型语言模型的训练特性,使其尽可能遵守用户请求。攻击者可以通过将有害指令嵌入到助手被指示总结的电子邮件或网页中,利用这一倾向。由于大型语言模型无法可靠地区分来自不可信方的邮件内容与直接输入到提示中的用户指令,过度顺从的LLM会忠实地遵循这些指令。迄今为止,Grok和其他大型语言模型的唯一对策是创建护栏,标记可疑指令并禁止执行。
Read full article
阅读全文
Comments
评论
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力