RAG incremental indexing: 5 rules for fresh answers

Analytics By Hai Ninh

Cover Image

RAG incremental indexing: 5 rules for fresh answers

Introduction

Someone updates the refund policy. Your assistant still quotes the old deadline. The source document is correct, the model follows its instructions, and the answer is wrong because retrieval supplied an outdated chunk.

RAG incremental indexing keeps a retrieval index aligned with source changes without rebuilding every document on every run. That requires more than detecting new files. Updates must replace the right chunks, deletes must stop old content from reaching answers, and retries must not restore an earlier version.

Start with a scheduled reconciliation job if your freshness requirements allow it. Add event-driven ingestion when waiting for the next scan is too slow. Either way, keep a durable record of what each document should contain and what the index currently serves.

The five rules below describe an implementation pattern, not a packaged framework or a benchmark. The first decision is what counts as the same document.

1. Give documents stable identities and separate versions

A filename can change. A URL can redirect. Neither is automatically a reliable primary key. Prefer an immutable source identifier, scoped to its tenant and source system. Keep the current URL separately so citations can follow a rename without creating a second document.

Track these fields in a durable document manifest:

Field

What it tells you

Tenant, source, and document ID

Which source object owns these chunks

Source revision

Whether an incoming update is newer than the accepted one

Content hash

Whether the extracted content changed

Pipeline fingerprint

Which parser, chunker, and embedding configuration produced the index

Active generation

Which completed chunk set retrieval may use

Chunk IDs or generation lookup

How to find records for verification and cleanup

Deleted state and processing status

Whether the document may be served and whether work remains

A content hash detects equality; it does not establish ordering. If revision 12 arrives after revision 13, different hashes alone cannot tell the worker which is newer. Use the source's revision semantics or an ordered change position. Opaque version tokens must not be compared as ordinary numbers or alphabetically.

Hash the representation you actually index. Normalizing line endings may be appropriate; deleting punctuation or flattening tables can hide meaningful changes. Track permissions and citation metadata separately if they are not part of that representation. An unchanged text hash must not prevent an access-policy update.

Also version the processing configuration. The same source text can produce different chunks after a parser change, and a different embedding model can require rebuilding vectors even when no document changed.

Think of the manifest as a library catalog: it identifies the current edition and where its pages belong. A pile of pages with similar text cannot replace that catalog. Once identity is reliable, replacement becomes a manageable operation.

2. Make RAG incremental indexing a staged replacement

Deleting all old chunks before creating new ones leaves a gap if embedding fails. Writing new chunks beside old ones creates a different problem: retrieval may return both versions.

One approach is to give each replacement a generation identifier and separate preparation from activation:

  1. Read the source revision and the manifest's current state.

  2. Build the replacement chunks under a new generation, using reproducible IDs for retries of that generation.

  3. Write the chunks and verify the expected set, metadata, and query visibility.

  4. Activate the generation only if it is still the intended source revision and no newer update or delete has superseded it.

  5. Retire the previous generation and clean up unreferenced records after any required retention period.

The activation step needs concurrency control. A transaction or compare-and-set in the manifest can reject a worker whose expected state is outdated. Where workers overlap, use ownership and fencing appropriate to your storage system so an expired worker cannot later commit stale work.

This pattern also needs a retrieval-side contract. Updating a manifest does not magically hide old vectors in another database. Retrieval must enforce the active generation and current permissions before any chunk enters the model context. Where a vector database cannot filter directly against that manifest, applications can validate candidates against it, but must account for stale candidates crowding out valid results. Additional fetching, cleanup, or a different indexing layout may be needed to maintain recall.

For a small system, a serialized per-document replacement with a documented temporary availability gap may be simpler. Choose that trade-off explicitly rather than describing a multi-request delete-and-insert sequence as atomic.

Prepare and validate replacement chunks, activate only the current complete generation, then retire old chunks. Retrieval enforces the active generation and current permissions.

The diagram separates writing chunks from making them eligible for answers. The retrieval check is part of the design, not an automatic consequence of updating the manifest.

The staged approach makes failure recovery easier to reason about, but only if a retry addresses the same records.

3. Make retries repeatable without accepting stale writes

Use stable point IDs for repeated writes of the same document generation. Random IDs on every retry can accumulate duplicate chunks. Include tenant, source identity, generation, and chunk position in the ID input so different documents do not collide accidentally.

This standalone example creates UUID-shaped IDs from a caller-supplied generation. It uses only Python's standard library: See the official documentation for these implementation details.

# Build repeatable point IDs for chunks in one document generation.
import json
import uuid


def chunk_id(tenant, source, document, generation, position):
    identity = json.dumps(
        [tenant, source, document, generation, position],
        ensure_ascii=False,
        separators=(",", ":"),
    )
    return str(uuid.uuid5(uuid.NAMESPACE_URL, identity))


first = chunk_id("tenant-a", "helpdesk", "policy-42", "revision-13-build-2", 0)
assert first == chunk_id("tenant-a", "helpdesk", "policy-42", "revision-13-build-2", 0)
assert first != chunk_id("tenant-a", "helpdesk", "policy-42", "revision-14-build-2", 0)
assert first != chunk_id("tenant-b", "helpdesk", "policy-42", "revision-13-build-2", 0)

This is an identity helper, not a synchronization engine or a security boundary. UUID generation does not validate revisions, authorize access, or make several writes transactional. Assign a different generation when the source revision or processing configuration changes. If rebuilding a generation can yield different chunk sets, persist its build plan or allocate a new generation rather than silently changing an existing one.

Database write semantics still matter. The Qdrant upsert reference states that a point with an existing ID is overwritten. Repeating an identical write can therefore address the same point, but an old payload can overwrite a newer one if your application reuses that ID incorrectly. Upsert is not an ordering guarantee.

Persist enough progress to recover after a crash between vector writes and manifest activation. A retry can verify staged records and complete activation, or discard an obsolete generation. Retain the work item until its required outcome is durable; a successful HTTP response is not a substitute for your database's documented completion and visibility semantics.

Updates are only half the lifecycle. A document that disappears needs an equally deliberate path.

4. Treat deletes and permission changes as serving decisions

A deleted source document can remain in vector storage indefinitely unless your ingestion process notices and handles deletion. Re-embedding changed documents does not remove absent ones.

On an authoritative delete, mark the document unavailable in the manifest and enforce that state before constructing model context. Then remove its vectors and other derived copies according to your retention policy. Keep a versioned deletion record long enough to reject delayed older updates; otherwise a retried ingestion job can resurrect the document.

Soft deletion is a serving control, not physical erasure. If you need to remove retained data, consider chunk text, vector payloads, extracted files, caches, logs, and backups as well as the primary index. Define what deletion means for each store instead of calling a single metadata flag complete erasure.

Permission changes deserve the same attention even when the document text is unchanged. Check the requester's current access before including retrieved content in a prompt. Invalidate or revalidate cached results where access changed. Do not rely on the language model to withhold content it should never have received.

Scheduled scans need a careful absence rule. A failed page request, expired credential, or incomplete source listing is not proof that every missing document was deleted. Compare absence only against a successfully completed inventory with the expected scope, and distinguish source deletion from loss of ingestion access.

Change streams can make detection faster. The Debezium PostgreSQL connector documentation describes database change events, including delete handling. A database change stream still needs an application mapping from changed rows to affected documents; a row event is not automatically a complete RAG indexing task.

Use events to reduce delay and reconciliation to detect drift. With the lifecycle covered, check how you will know it is working.

5. Measure freshness and rehearse a full rebuild

Define freshness at the serving boundary: how long after a source change can a retrieval request obtain the new authorized content? Track deletion propagation separately, since an obsolete paragraph and an exposed restricted document have different consequences.

Useful operational signals include the oldest unprocessed change, failed document count, source revision versus active revision, and staged generations that never became active. A queue length alone cannot distinguish a healthy burst from one document stuck for days.

Test the failure cases that threaten your design:

Scenario

Expected behavior

The same update arrives twice

One logical active generation, without duplicate accumulation

An older update arrives after a newer one

The older update cannot become active

Embedding fails halfway through a document

An incomplete generation stays unavailable

A worker crashes after vector writes

A retry can reconcile staged records and manifest state

A delete arrives while an update runs

The delayed update cannot restore the deleted document

An inventory scan stops halfway

Unseen documents are not bulk-deleted as if absence were confirmed

Access changes without text changes

Retrieval enforces the new policy

These are acceptance criteria to implement and test, not results claimed for the example helper.

An embedding-model or chunking migration often warrants a separate index generation. Build it from a defined source snapshot, apply subsequent changes including deletes, and compare retrieval quality before switching traffic. Keep the old index only as long as your rollback and retention policies permit. If you need rollback after new writes, plan how both versions stay current.

The Qdrant collection aliases documentation describes building a second collection and changing aliases atomically. That feature covers the collection reference change. Your application must still pair the collection with the correct query embedding model and retrieval configuration; an alias switch does not update those automatically. Matching vector dimensions alone does not establish embedding compatibility.

Before adopting any migration procedure, trace one document through both the normal update path and its failure cases.

Start with one document you can trace

RAG incremental indexing is easier to operate when you can explain one document's state without searching through unrelated worker logs. Find its source revision, identify the active chunks, and show why a delayed job cannot replace them with older content.

Begin with stable identities, a durable manifest, and a reconciliation schedule that meets your needs. Add event-driven ingestion when the delay matters. Before tuning throughput, run an update, a failed replacement, a delete, and a permission change through the same document and inspect what retrieval can actually return.

You do not need the most elaborate pipeline. You need to know which version the next answer will use, and have a dependable way to change it.

Author

Hai Ninh

Author

Hai Ninh

Software Engineer

Love the simply thing and trending tek

More to read

Related posts