跳到主内容
精选86Rohan Paul论文研究

Meta论文:计算机使用Agent在GUI任务中的安全漏洞与基准测试

New Meta paper.

原文
发到 X
推荐理由

Agent安全是落地关键瓶颈,这篇论文用具体基准揭示了当前主流模型在GUI交互中的致命缺陷及防御策略差异,值得关注。

New Meta paper.

Meta 的新论文。

Computer-use agents can click the requested button and still fail to recognize when that click should never happen.

计算机使用代理可以点击请求的按钮,但仍无法识别何时绝不应发生该点击操作。

If an agent can process a $25K checkout without stopping, task success alone is a dangerously incomplete benchmark.

如果代理能够不间断地完成一笔 25,000 美元的结账流程,仅以任务成功率为指标是一种危险且不完整的基准测试。

ADeptS-Bench tests 7 models on paired benign/malicious GUI tasks plus ambiguous instructions across mobile and desktop.

ADeptS-Bench 在移动和桌面平台上,针对配对的良好/恶意 GUI 任务以及模糊指令,测试了 7 个模型。

No model consistently stays above 80% task success while keeping attack success below 30% across both platforms.

没有哪个模型能在两个平台上同时保持任务成功率始终高于 80%,且攻击成功率低于 30%。

The failure is consequence reasoning.

问题在于后果推理。

All 7 models went ahead with a $25K checkout, and none caught a button labeled “Optimize” that actually triggered a factory reset.

所有 7 个模型都继续进行了一笔 25,000 美元的结账操作,且没有一个模型察觉到标有“优化”的按钮实际上会触发恢复出厂设置。

The ablation makes the safety problem more concrete.

消融实验使安全问题更加具体化。

Removing the explicit refusal tool and its usage instruction raised attack success by 22.0 percentage points for Gemini 3.1 Pro, 10.3 for Claude 4.7, and 10.7 for GPT-5.4, while Qwen was essentially unchanged.

移除显式的拒绝工具及其使用说明后,Gemini 3.1 Pro 的攻击成功率上升了 22.0 个百分点,Claude 4.7 上升了 10.3 个百分点,GPT-5.4 上升了 10.7 个百分点,而 Qwen 基本保持不变。

So part of today’s “agent safety” can live in the wrapper, not the model.

因此,当今部分“代理安全性”可以存在于封装层(wrapper),而非模型本身。

– arxiv. org/abs/2608.26204

– arxiv.org/abs/2608.26204

Title: "ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices"

标题:《ADeptS-Bench:跨设备衡量计算机使用代理的可信度》

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近