PII redaction for RAG: 5 checks before embedding

Automation By Hai Ninh

Cover Image

PII redaction for RAG: 5 checks before embedding

A support document contains a customer's account number. Your assistant hides it in the final answer, but the retrieval system still stores the original paragraph as chunk text. The answer looks clean; the stored data has not changed.

PII redaction for RAG starts before embedding when the retrieval task does not need those identifiers. Detect and replace unnecessary personal information before creating retrieval chunks, then restrict which processed documents can reach the embedder. Keep access checks and response filtering too: automated detection can miss information, and replacing names does not make a document anonymous.

This guide covers five design checks for that ingestion path, including document boundaries, replacement choices, failure handling, and existing indexes. It describes an engineering pattern, not a tested deployment or a legal determination. Start by deciding what information the system actually needs.

1. Define what PII redaction for RAG should remove

A general policy assistant usually needs the refund rules, not the identity of the customer whose support ticket supplied an example. An authenticated order-status assistant may need a customer-specific lookup. Those systems require different data policies. Removing every identifier indiscriminately can make the second system unusable without making its authorization problem disappear.

List the personal-data categories the retrieval task needs, those it should remove, and documents that should never enter this index. Include filenames, URLs, table headers, and metadata, not only paragraph text. A redacted body can still travel with a revealing title or citation link.

Do not treat an embedding as encryption. The peer-reviewed paper Text Embeddings Reveal (Almost) As Much As Text demonstrates text recovery from embeddings under the authors' experimental conditions. That does not mean every stored vector contains a readable account number or that every embedding model exposes identical information. It does mean you should not assume that transforming text into vectors removes its privacy implications.

Stored chunk text is a more direct exposure path. So are raw-text logs, queue messages, and caches if your architecture creates them. Masking the answer does not erase those copies, although it can reduce what a reader sees.

Think of intake screening at a records office: checking a document before photocopying reduces what spreads to later copies. It does not prove the screening caught everything. With that distinction clear, choose where screening belongs.

2. Detect with enough context before retrieval chunking

Apply detection to the extracted document before splitting it into retrieval-sized chunks where practical. Splitting first can divide an identifier or separate a value from the label that explains it. A table cell containing a number means more when the detector can also see the account-number column header.

This is an ordering recommendation, not a requirement to send an unlimited document into one model request. Large documents may exceed detector limits. Use bounded sections with sufficient overlap and preserved structural context, reconcile overlapping findings, and account for every section before releasing the processed document. Do not silently truncate the input.

For PDFs, inspect extraction quality first. If OCR drops digits or a parser detaches table headings, the detector receives damaged evidence. Route unreadable pages and unsupported formats to a restricted review or retry path rather than treating empty output as proof that no PII exists.

The property you can enforce is narrower than "the index contains no PII": every released document passed the configured detection, replacement, and validation policy, and downstream ingestion uses that processed artifact. Detection misses remain possible.

Operational failures should block release. Define a policy for ambiguous findings separately: detector confidence scores are not universal measures of safety. Route cases outside that policy to review. Once the input boundaries are accounted for, the next question is what replaces the detected values.

3. Choose replacements without promising anonymity

Pattern matching and model-based recognition solve different parts of detection. Rules and validators can find structured formats such as emails or identifiers with checksums. Named-entity models can use context to recognize names and addresses. Neither is sufficient evidence of complete coverage, and each needs evaluation on the languages and document types you ingest.

Presidio, originally developed at Microsoft and now transitioning to community ownership under Data Privacy Stack, provides separate analysis and anonymization components. Evaluate its recognizers, supported languages and replacement operators against your corpus; it does not guarantee complete detection.

Replacement choices affect both privacy and retrieval quality:

Approach

Example replacement

Main trade-off

Remove or mask the detected value

[REDACTED]

Loses the value and may remove useful relationships

Use an entity-type placeholder

[PERSON]

Preserves the category but can merge distinct people in the reader's view

Use a scoped entity token

[PERSON_1]

Can preserve repeated references if entity matching is correct; introduces linkability

Use a synthetic substitute

A generated replacement name

May preserve readable prose but can be mistaken for a factual identity

These are engineering choices, not mutually exclusive legal categories. Replacing identifiers with tokens can be a form of pseudonymization. A token does not automatically remain consistent: entity matching, token assignment, and retry behavior must provide that consistency.

Choose the narrowest consistency scope the task needs. Document-scoped tokens can preserve references inside one document without deliberately linking a person across the whole corpus. Global tokens enable cross-document joins, but also make linkage easier. Do not reuse identity mappings across tenants.

If restoration is necessary, protect the mapping separately from the vector store and exclude it from routine retrieval. Authorize any restoration against the original record and current requester; do not blindly replace every token emitted by a language model. A model can produce a token that the current answer should not be allowed to resolve.

Even without a mapping, surrounding facts can identify someone. Test those residual clues alongside the replacement behavior before connecting the embedder.

4. Enforce the processing boundary and test failures

Build the ingestion interface around a processed artifact, not a status label attached to raw text. The worker that prepares embedding requests should load the released revision of that artifact and carry its document and policy identifiers into downstream records. Restrict who can mark an artifact released or change it afterward.

An audit record can include the source revision, detector and policy versions, processing outcome, and a reference or digest for the processed artifact. Avoid storing detected values in ordinary logs. Treat excerpts, entity metadata, and digests as potentially sensitive too.

A content hash alone does not prove that redaction succeeded or that an external service embedded the intended text. It supports traceability only when the recorded artifact is tied to the actual request-building path. Record chunk-level provenance if that path splits or otherwise transforms the document.

Raw documents pass through detection and replacement, policy validation and release, then chunking, embedding and storage. Failed processing is withheld; released content can still contain missed personal information.

The diagram shows the permitted ingestion path. The combined final stage keeps chunking, embedding, and storage in their execution order; it does not imply one atomic database operation. Failed or incomplete processing stays outside that path.

Use a labeled test set to evaluate detection by entity category and document type. Measure missed entities and false positives, then check whether the processed content still supports correct answers. The following are acceptance criteria to implement, not test results claimed for this article:

Failure case

Expected behavior

Detector timeout or missing page

No release to embedding; restricted retry or review

Identifier split across processing sections

Boundary handling is evaluated with labeled examples

Raw title remains beside a processed body

Metadata policy catches the field before release

Retry after a policy change

Artifact and policy versions cannot be confused

Two people share the same name

Token assignment does not assume they are one entity

A missed identifier survives detection

Access restrictions still apply; the finding triggers correction

Sampling can reveal failures, but cannot establish that every document is clean. Your release policy needs both measurable quality criteria and a response when a miss is found. That response also has to cover material indexed before this pipeline existed.

5. Cover retained copies, retrieval, and existing indexes

Adding redaction to new ingestion does not repair an old index. Inventory existing chunk text, vectors, source copies, caches, exports, and backups according to what your system actually stores. If you discover sensitive data that should not be served, restrict access while assessing it rather than waiting for a scheduled rebuild to finish.

Reprocessing creates a new set of derived records. Plan how to stop serving the superseded chunks, remove or expire retained copies under the applicable policy, and prevent an old retry from recreating them. Maintain source-to-derived-record lineage so correction and deletion requests can reach the affected stores.

At query time, enforce current document permissions before any retrieved content enters the model context. Permission checks do not automatically detect old unredacted documents: eligibility needs to reflect the processing policy too. Output filtering provides another chance to catch sensitive content, but it is also fallible. Neither control replaces ingestion processing.

The ICO's pseudonymisation guidance explains that pseudonymized personal data remains within data-protection law. Do not label a tokenized index anonymous or compliant merely because direct identifiers were replaced. Determine the applicable obligations with qualified privacy or legal support, including purpose, access, retention, and any external processing arrangements.

The detection service itself may receive raw personal data. Check where it runs and how its provider handles requests before sending documents to it. Local operation changes that exposure boundary but does not remove your own storage and access responsibilities.

These controls belong in one document lifecycle, so start by tracing a small representative set through the complete path.

Start with a document you can inspect

PII redaction for RAG is useful when it reduces unnecessary personal information before that information spreads into retrieval storage. Its value depends on the data policy, extraction quality, detection coverage, and whether the embedder actually receives the released text.

Choose representative documents, including awkward tables and boundary cases. Compare the source, detected spans, replacements, embedding inputs, and retrieval answers. Keep sensitive samples in an appropriately restricted environment and record failures without copying their raw values into routine logs.

Use what you find to set the release criteria and size any reprocessing work. The goal is a pipeline whose decisions you can inspect and correct, with no assumption that a clean-looking answer means the stored data is safe.

Author

Hai Ninh

Author

Hai Ninh

Software Engineer

Love the simply thing and trending tek

More to read

Related posts