RAG投毒可致模型过度自信,注意力坍缩可作检测信号
RAG poisoning can make a model more confident, which is exactly why confidence-b…
RAG poisoning can make a model more confident, which is exactly why confidence-based detectors can fail.
RAG投毒可以使模型更加自信,这正是基于置信度的检测器可能失效的原因。
This paper finds malicious retrieved documents can increase token confidence and output consistency.
本文发现,恶意检索文档可以增加令牌置信度和输出一致性。
So uncertainty-based detectors can miss the attack because poisoning can create false confidence.
因此,基于不确定性的检测器可能漏掉攻击,因为投毒可以制造虚假的自信。
Under attack, attention becomes concentrated on poisoned documents instead of staying spread across the retrieved evidence.
在攻击下,注意力会集中在被投毒的文档上,而不是分散在检索到的证据中。
The authors call this Attention Collapse.
作者将这种现象称为“注意力崩溃”。
So for RAG security, checking only the final answer or its confidence may miss the warning.
因此,对于RAG安全而言,仅检查最终答案或其置信度可能会错过警告信号。
Monitoring how attention is distributed across retrieved documents could expose poisoning before the answer visibly breaks.
监控注意力在检索文档中的分布,可以在答案明显出错之前暴露投毒行为。
– arxiv. org/abs/2608.06947
– arxiv.org/abs/2608.06947
Title: "When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse"
标题:“当上下文咬人:通过文档级注意力崩溃检测RAG投毒”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力