← Files RAG & GenAI CopilotARCHIVED FILE
skills/production-rag-genai-copilot/references/rag_security_guardrails.md
2.74 KB · Oct 4, 2026 · 12:36 UTC
# RAG Security and Guardrails ## Purpose Use this file to review security boundaries in retrieval and grounded-generation systems. Retrieved content is untrusted data, not instructions. # 1. Prompt injection Threats include: - direct user injection; - indirect instructions in retrieved/web content; - hidden/multimodal instructions; - malicious metadata; - poisoned knowledge-base documents. RAG does not eliminate prompt injection risk. # 2. Authorization Enforce access before or during retrieval. Do not rely on: - prompt instructions; - post-generation redaction; - “the model should not mention it.” Use authoritative user/tenant/resource filters at the retrieval/data boundary. # 3. Cross-tenant leakage Check: - tenant metadata; - filter construction; - cache keys; - shared indexes; - query rewrite; - reranking; - context assembly; - trace/log data. One missing filter can make ranking quality irrelevant. # 4. RAG poisoning Threat model: an attacker changes/indexes content designed to manipulate retrieval or generation. Controls: - source allowlists; - trusted ingestion paths; - provenance; - content ownership; - document versioning; - anomaly/review workflow; - separation of instructions from retrieved data. # 5. Secrets and PII Do not index or log secrets unnecessarily. Define: - redaction; - retention; - encryption boundary; - access control; - trace policy; - deletion/update propagation. If sensitive data must be searchable, protect it at the same level as the source system. # 6. Unsafe URLs/files When retrieval includes external links/files: - validate source; - sandbox parsing/execution where relevant; - block dangerous file behavior; - do not execute retrieved instructions/code automatically. # 7. Exfiltration Consider outputs that can leak data through: - links/URLs; - markdown/image requests; - tool calls; - citations; - logs; - third-party APIs. Restrict outbound destinations when needed. # 8. Tool-enabled RAG If retrieved content can influence tools/actions: - keep tool permissions narrow; - validate parameters; - require approval for consequential actions; - do not treat retrieved text as authority to act. # 9. Security evals Include: - direct injection; - indirect injection; - cross-tenant query; - poisoned document; - secret-seeking query; - malicious URL/file; - unauthorized citation/source. Measure both blocked attacks and false positives. # 10. Incident response For suspected leakage: 1. contain access; 2. preserve traces/evidence; 3. identify affected tenants/documents; 4. revoke/rotate credentials if exposed; 5. remove poisoned content safely; 6. rebuild/reindex only with scoped recovery plan; 7. verify with targeted security evals. Do not casually purge/rebuild everything before preserving evidence.
SHA-256: fb978b646315fc3401c222737e5f8ef6129dac40de9f0a6b0a646e8532298d25