Skip to content

Reference

RAG, end to end

Retrieval-augmented generation is where an attacker’s document meets your model’s trust. RAG is spread across the book because it touches every layer - the context window that ingests it, the data layer that stores it, the method that tests it. This page pulls it into one place: what the pipeline is, where it breaks, how to hold it, and which chapter carries each piece in depth.

What RAG is

RAG lets a model answer from your documents without retraining. The pipeline is six steps: ingest, chunk, embed, store, retrieve, ground. A question is turned into a vector, the closest chunks are pulled from a vector database, and those chunks are pasted into the context before the model answers. The concrete request-to-answer trace is in Start here - the basics.

The one fact that makes it a security topic: the retrieved text lands in the same context as your instructions. The model has no reliable way to tell a trusted system instruction from a sentence in a document it just fetched. So every document that can be indexed is a candidate instruction, and the vector store holds a derived copy of whatever it was built from.

The attack surface

AttackWhat it isDepth
Knowledge-base poisoningGet a malicious instruction indexed, so the model retrieves and obeys it laterVI.4 · II.2
Retrieval manipulationCraft content to win the similarity match for a target query (the PoisonedRAG line - a few passages can control answers)VI.4
Authority spoofing (DACSI)Non-imperative, metadata-like text impersonating a provenance or policy signal, evading imperative-injection filtersII.2
Indirect injection at scaleInstructions planted in non-rendered fields of pages a crawler indexes, never seen by a humanII.2
Cross-tenant leakageA shared multi-tenant store with no role-aware retrieval surfaces another tenant’s documentsV.3
Embedding inversionReconstruct source text from stored vectors - storing embeddings is not anonymizationV.3
Membership inferenceDecide whether a specific record is in the store or training set, from similarity signalsI.3

How you defend it

The retrieved chunk is untrusted input that happens to look authoritative. The defenses treat it that way:

  • Role-aware, entitlement-checked retrieval. Re-check the requesting user’s permissions at query time, against the documents being returned - not just at ingest. This is the single control that closes cross-tenant leakage. See V.3 · The data layer.
  • Spotlight and delimit retrieved text. Mark retrieved content as data, never instruction, and keep it out of the instruction channel. The mechanics are in II.2 · Prompt injection & the LLM attack surface and II.5 · Guardrails - what holds, and how to prove it.
  • Provenance on every chunk. Validate and sign ingested sources; tag each chunk with where it came from, and distrust chunks whose source is untrusted.
  • Secure the vector store as raw data. Encrypt vectors at rest, lock down the store, and minimize what you embed - it holds the data it was derived from (V.3 · The data layer).
  • Gate the ingest. Treat the corpus as an attack surface: scan your own crawl or knowledge base for instruction-like text in non-rendered fields before it is indexed.

Where it is covered in depth