跳到主内容
@wquguru
精选90r/Rag(Reddit)技巧与观点

构建权限感知RAG:在检索层拦截越权访问的完整工程实践

I built a RAG service where the LLM never sees documents the user isn't allowed to read (permission-aware, multi-tenant)

原文
发到 X
推荐理由

做企业级RAG的同学必看,这篇给出了从元数据设计到对抗评估的完整安全落地方案,直接解决多租户数据隔离痛点。

Most RAG tutorials end at "embed docs, retrieve top-k, stuff into prompt." That works until you have more than one user. Then you hit the question nobody covers: what stops Alice from asking a question and getting an answer built from Bob's confidential docs?

大多数 RAG 教程在“嵌入文档、检索 top-k、塞入提示词”这一步就停止了。这在只有一个用户时行得通,但当你面对多个用户时,就会遇到一个无人解答的问题:是什么阻止了 Alice 提问并获得基于 Bob 机密文档构建的答案?

Prompt instructions like "don't reveal other tenants' data" aren't a security boundary. Once a restricted chunk is in the context window, you've already lost.

诸如“不要泄露其他租户的数据”之类的提示指令并不是安全边界。一旦受限的文本块进入了上下文窗口,你就已经输了。

So I built GateKeep RAG, where the access check happens before the LLM is involved:

因此我构建了 GateKeep RAG,其访问检查发生在 LLM 介入之前:

  • Access metadata on every chunk. Each chunk carries its tenant ID, required roles, and clearance level, set at ingestion.
  • Filtering at retrieval time. The query is pre-filtered by the caller's tenant and role directly in the vector database query, so restricted chunks never leave the database or reach the prompt. A second relational check in PostgreSQL acts as defense-in-depth before prompt assembly.
  • Tenant isolation by design. Cross-tenant leakage is prevented structurally, not by asking the model nicely.
  • Tamper-evident audit log. Every query records what was retrieved, for whom, and when. Each entry is written to an append-only, tenant-scoped SHA-256 cryptographic hash chain locked with PostgreSQL advisory transactions (pg_advisory_xact_lock), so any retroactive row modification or deletion breaks the chain and is detected by /v1/audit/verify.
  • Adversarial eval harness. 440 hand-written queries and 4,972 principal-query counterfactual pairs tested against a live stack. We injected 128-bit CSPRNG canary tokens into all 210 corpus chunks: across 823,996 checks (inspecting prompts, answers, citations, and raw JSON payloads), leak rate was 0.00%. We also verified counterfactual invariance—unauthorized users get the exact same refusal whether restricted docs exist in the DB or are physically deleted.
  • 在每个文本块上访问元数据。每个文本块在摄入时都携带其租户 ID、所需角色和清除级别。
  • 检索时的过滤。查询直接在向量数据库查询中按调用者的租户和角色进行预过滤,因此受限文本块永远不会离开数据库或到达提示词。PostgreSQL 中的第二个关系型检查作为提示词组装前的纵深防御。
  • 按设计实现租户隔离。跨租户泄露是通过结构性方式防止的,而不是靠礼貌地请求模型。
  • 防篡改审计日志。每次查询都会记录检索了什么、为谁检索以及何时检索。每条条目都写入仅追加的、租户范围的 SHA-256 密码学哈希链,并使用 PostgreSQL advisory 事务(pg_advisory_xact_lock)锁定,因此任何回溯性的行修改或删除都会破坏链条并被 /v1/audit/verify 检测到。
  • 对抗性评估框架。使用 440 个手写查询和 4,972 个主体-查询反事实对针对实时堆栈进行测试。我们将 128 位 CSPRNG 哨兵令牌注入到所有 210 个语料库文本块中:在 823,996 次检查(检查提示词、答案、引用和原始 JSON 负载)中,泄露率为 0.00%。我们还验证了反事实不变性——无论受限文档存在于数据库中还是被物理删除,未经授权的用户都会得到完全相同的拒绝响应。

Stack:

技术栈:

  • Backend: FastAPI (Python 3.11), SQLAlchemy, Alembic
  • Databases: PostgreSQL 16 (relational & audit) + Qdrant (vector search with payload pre-filtering)
  • Embeddings: all-MiniLM-L6-v2 (SentenceTransformers, 384d)
  • LLM: Ollama (llama3.2:3b) / pluggable local or cloud model
  • Frontend: React + TypeScript + Vite + Tailwind CSS
  • 后端:FastAPI (Python 3.11), SQLAlchemy, Alembic
  • 数据库:PostgreSQL 16 (关系型与审计) + Qdrant (带有效载荷预过滤的向量搜索)
  • 嵌入模型:all-MiniLM-L6-v2 (SentenceTransformers, 384d)
  • LLM:Ollama (llama3.2:3b) / 可插拔本地或云模型
  • 前端:React + TypeScript + Vite + Tailwind CSS

Repo: https://github.com/saturn-16/GateKeep-RAG

仓库:https://github.com/saturn-16/GateKeep-RAG

What I learned:

我的心得:

  • Filter in the retrieval layer, not the prompt. It's the only version you can mathematically and structurally reason about.
  • Evals for leakage are different from evals for quality. Quality is recall/MRR; security is adversarial counterfactuals and canary tokens. You need both, and security tests should pass even if you set similarity threshold to 0.00.
  • Similarity thresholds cannot reliably reject unanswerable questions. Dense-only retrieval with cosine score thresholds struggles with false positives—if a user asks about dental plans and none exist, general medical docs still score above a 0.35 threshold on loose semantic overlap. Cosine cutoffs alone won't reject unanswerable queries; that requires candidate-scoped hybrid search (BM25 + dense) or explicit retrieval intent gates.
  • 在检索层进行过滤,而不是在提示词中。这是唯一可以从数学和结构上进行推理的版本。
  • 泄露评估与质量评估不同。质量评估关注召回率/MRR;安全评估则涉及对抗性反事实样本和金丝雀令牌。两者都需要,且即使将相似度阈值设为 0.00,安全测试也应通过。
  • 相似度阈值无法可靠地拒答不可回答的问题。仅使用余弦分数阈值的稠密检索在处理假阳性时存在困难——如果用户询问牙科保险计划而系统中不存在相关文档,一般医疗文档仍可能因宽松的语义重叠而在 0.35 的阈值下得分较高。仅靠余弦截断无法拒答不可回答的查询;这需要候选范围混合搜索(BM25 + 稠密)或显式的检索意图门控。

I'm a CSE (cybersecurity) student, so this started as a "how would I break this?" project. I'd love feedback, especially on attack scenarios I haven't thought of (prompt injection via documents, metadata spoofing, etc.).

我是一名网络安全(CSE)专业的学生,因此这个项目最初是出于“我会如何攻破它”的想法。我很期待反馈,特别是关于我未曾想到的攻击场景(例如通过文档进行的提示注入、元数据伪造等)。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件