Skip to main content

Command Palette

Search for a command to run...

Deleting Personal Data from a RAG System in .NET

Following an erasure request through documents, chunks, vectors, caches and background jobs

Updated
•21 min read•View as Markdown
Deleting Personal Data from a RAG System in .NET
P
Senior Software Engineer specialising in cloud architecture, distributed systems, and modern .NET development, with over two decades of experience designing and delivering enterprise platforms in financial, insurance, and high-scale commercial environments. My focus is on building systems that are reliable, scalable, and maintainable over the long term. I’ve led modernisation initiatives moving legacy platforms to cloud-native Azure architectures, designed high-throughput streaming solutions to eliminate performance bottlenecks, and implemented secure microservices environments using container-based deployment models and event-driven integration patterns. From an architecture perspective, I have strong practical experience applying approaches such as Vertical Slice Architecture, Domain-Driven Design, Clean Architecture, and Hexagonal Architecture. I’m particularly interested in modular system design that balances delivery speed with long-term sustainability, and I enjoy solving complex problems involving distributed workflows, performance optimisation, and system reliability. I enjoy mentoring engineers, contributing to architectural decisions, and helping teams simplify complex systems into clear, maintainable designs. I’m always open to connecting with other engineers, architects, and technology leaders working on modern cloud and distributed system challenges.

A customer asks you to remove their personal data. You delete their database record, remove the uploaded PDF and close the request. Later that afternoon, your AI assistant answers a question using a paragraph from the same PDF. The customer's name, address and circumstances are still sitting in your search index.

You remove those search documents too. Overnight, an ingestion job retries and writes them back. This is a realistic failure in a retrieval augmented generation system. Ingestion creates multiple representations of the same information. Retrieval feeds that information into more systems, and generated answers can become persistent records themselves. Deletion has to follow those transformations and survive the asynchronous work happening around them.

In this article, we’ll design a deletion workflow for a .NET application using a relational database, Blob Storage, a vector enabled search index and background workers. Azure AI Search provides the concrete vector store example, but the design also applies to other stores once you account for their filtering, consistency and deletion behaviour. The objective is to remove the personal data covered by an assessed erasure request, propagate the change through its derived artefacts, prevent obsolete work from reintroducing it and collect evidence of the result. That requires more than a delete endpoint.

Decide what the request covers

The right to erasure under GDPR Article 17 applies in specified circumstances and has exceptions. A request can affect some processing while other records must remain for a lawful purpose, such as a legal obligation or the establishment, exercise or defence of legal claims. The Irish Data Protection Commission explains those conditions and exceptions in its guidance on the right to erasure. Your application should therefore receive an assessed scope. Identity verification, the grounds for erasure and any retention decisions belong in the request handling process. An engineering workflow then carries out that decision consistently.

Imagine a document contains details about three people. One requests erasure. Depending on the assessed scope, you might remove the whole document, create a properly redacted replacement, or retain a restricted original for a specific obligation while removing it from ordinary AI retrieval. Each outcome needs explicit handling. Retaining an original for one purpose does not automatically justify keeping it available to a general purpose assistant. Treat redaction as a new document version. A black rectangle in a PDF is insufficient if the underlying text, OCR layer, annotations or metadata still expose the information. Once an approved replacement exists, regenerate the text, chunks and embeddings from that replacement and remove the superseded derivatives. The engineering examples below assume that the scope has already been approved. They provide a way to execute and demonstrate the change, rather than deciding which data a particular organisation is legally entitled to retain.

Follow the information through the system

A source file rarely stays in one place. An email attachment becomes a blob, its extracted text becomes another blob, its paragraphs become search documents, and its embeddings become vector fields. Retrieved passages may then appear in a conversation, a cached answer, an exported report or an evaluation fixture.

The arrows describe provenance - one artefact helped create another. They are also the routes along which deletion or replacement must propagate. An embedding does not become harmless simply because it is an array of floating point values. For deletion planning, treat embeddings derived from affected content as artefacts in scope for assessment. Do not assume that a transformation makes information anonymous. In a typical RAG index, the vector also sits alongside readable chunk text and metadata, which may identify the person directly.

Deleting the vector field while keeping the chunk leaves an obvious copy behind. Deleting the text while retaining a vector and identifying metadata can leave a derived representation that still needs assessment. Removing the entire obsolete chunk record is often the simpler operational choice. You also need to include data outside the main ingestion path. Raw prompt logging, OCR responses, document previews, temporary files, failed message payloads and tracing exports can all hold personal data. A vector database is one participant in the workflow.

Build lineage while you ingest

You cannot reliably discover every affected record by asking the vector database for passages similar to a person’s name. Retrieval deliberately returns a limited set of relevant matches. It can miss spelling variations, identifiers, indirect descriptions, overlapping chunks and material that ranks poorly for that query. Use explicit identity and provenance instead. Give each source document a stable identifier and each version its own revision. Record which subjects have been associated with that source and which artefacts each processing step produces.

An application level chunk record could look like this:

public sealed record ChunkManifestEntry(
    Guid TenantId,
    Guid SourceDocumentId,
    long SourceRevision,
    string IndexName,
    string ChunkKey,
    Guid[] SubjectIds);

public sealed record DerivedArtifact(
    Guid ArtifactId,
    Guid TenantId,
    string Store,
    string Locator,
    Guid SourceDocumentId,
    long SourceRevision);

public sealed record ArtifactDependency(
    Guid ParentArtifactId,
    Guid ChildArtifactId);

These types describe the metadata, not a complete database schema. A relational implementation would normally use separate subject link and dependency tables, with indexes on tenant, subject and source identifiers. Store locators need restricted access because paths and names can themselves contain personal information. Source lineage and subject identification solve different problems. Lineage tells you where a document’s contents went. Subject identification tells you which source material is affected by a person’s request. Some documents arrive with authoritative customer identifiers; others contain free text requiring identity resolution and review. An entity detector can suggest links, but missed entities and ambiguous names still need a discovery process.

This is where legacy data makes the problem harder. If earlier ingestion created no lineage, plan an inventory and backfill. Reconcile known records against actual storage, inspect untracked stores, and use targeted searches or detection tools as additional discovery methods. A confident completion flag cannot compensate for an incomplete inventory.

Register outputs before making them usable

Consider a worker that writes a vector and then records its chunk key in SQL. If the process dies between those steps, the vector exists without a manifest entry. Your deletion workflow will never discover it through the manifest. Record the intended output before publishing it. Use deterministic artefact identities, persist a pending record and then perform the external write. If the worker crashes, reconciliation can look up the intended object and either finish registration or remove it.

Deterministic identities need to include the source revision and any chunking generation that changes the output set. A new splitter might produce twelve chunks where an older one produced eighteen. Overwriting twelve keys leaves six stale records unless the previous generation is explicitly cleaned up.

Generated outputs need provenance too. When an answer uses several passages, record the source artefacts that contributed context. You do not need to store an additional full copy of the prompt to establish that dependency. If personal information appears in the user’s question rather than retrieved context, that input also needs subject linkage and its own retention handling.

Where attribution is uncertain, an affected answer can be invalidated as a whole or routed for assessment. Automatically removing every descendant may be broader than the approved scope, but automatically preserving all generated text assumes away the possibility that it contains the same personal information. The manifest becomes a description of what exists, what is pending, what depends on what and what has already been removed.

Block affected content before deleting it

Deletion across several stores takes time. During that interval, the application should prevent the affected content from being used in the processing covered by the request. Keep an authoritative publication state for each source version. Retrieval resolves candidate chunks against that state before passing their text to the model. The same rule applies when serving cached answers, saved conversations or downloads. A cache hit must not bypass the current eligibility decision. This gate is particularly important when an external index still contains an old document while its deletion is being processed. The application can refuse to serve the content even though the external cleanup has not yet finished.

Keep the consistency requirement explicit. An asynchronously refreshed cache of allowed documents creates a window in which obsolete permissions remain effective. For the affected path, use a strongly consistent authority or a deliberately designed revocation mechanism. If the authority cannot be consulted, fail closed for that content. Requests already in progress require handling too. Revalidate before assembling context and before persisting or releasing the answer. Track affected active runs so they can be cancelled or their outputs discarded. Bytes already streamed to a client cannot be taken back; the request record should not pretend otherwise. For sensitive workflows, buffered output gives you a more useful final eligibility check than immediate streaming. A publication block is an interim control. The deletion workflow still has to remove the covered data from its stores.

Persist the request and propagation together

An endpoint that updates SQL and then publishes a message has a failure window. SQL can commit successfully while the message publish fails. The request appears accepted, but no worker ever receives it. Use a transactional outbox. In one local database transaction, persist the assessed operation, change the affected publication state and insert the outbox message. A dispatcher then publishes committed outbox records and retries when delivery fails. The HTTP response can acknowledge the operation with 202 Accepted and a status URL. Acceptance means that the request has been recorded; it does not mean every store has completed deletion.

The outbox delivers propagation intent reliably from that database. It does not create a distributed transaction across SQL, Blob Storage and the vector store. The coordinator handles partial progress explicitly.

A practical operation model separates a request from its store tasks:

public enum ErasureTaskState
{
    Pending,
    Running,
    RetryScheduled,
    AwaitingVerification,
    Verified,
    RequiresIntervention,
    RetainedByDecision
}

public sealed record ErasureTask(
    Guid OperationId,
    Guid TaskId,
    string Store,
    string Target,
    long SourceRevision,
    ErasureTaskState State);

Each task represents a specific target and action. Redaction, deletion and approved restricted retention should be distinguishable in the real schema. RetainedByDecision must reference an assessed decision; it must never become a convenient substitute for handling an error.

Make propagation observable and repeatable

Consumers should expect duplicate delivery. A worker might successfully delete an object and crash before acknowledging the message. The queue then delivers it again. Give each task a stable idempotency key, record attempts and use a durable claim or lease to coordinate workers. If the external operation succeeds but the local acknowledgement fails, retrying the same deletion should be safe. An inbox record can deduplicate message handling, but it cannot atomically wrap an unrelated provider’s API call. Dont mark a task complete before doing the work. Equally, do not assume that finding an inbox entry means the external deletion finished. Persist execution state and recover abandoned claims.

A store specific adapter needs separate execution and verification methods:

public interface IErasureStore
{
    string Name { get; }

    Task ExecuteAsync(
        ErasureTask task,
        CancellationToken stopToken);

    Task<ErasureVerification> VerifyAsync(
        ErasureTask task,
        CancellationToken stopToken);
}

public sealed record ErasureVerification(
    bool Verified,
    DateTimeOffset ObservedAt,
    string EvidenceCode);

These are application interfaces. Each adapter translates the target into the provider's API, classifies failures and supplies evidence appropriate to that store. A vector store adapter might verify exact keys and revision filters. A blob adapter might inspect the current object, snapshots and versions. A processor adapter might track a deletion request and its acknowledgement. Transient failures need bounded retries with backoff. Invalid credentials, malformed targets and unsupported retention policies need intervention. Failed tasks remain visible after messages enter a dead letter queue. Track operation age, verification age and unresolved stores so the request cannot disappear into background infrastructure. Azure Durable Functions can coordinate this long running work, or you can use queue workers backed by a persisted operation state. With Durable Functions, external calls belong in activities and orchestration must respect replay constraints. Keep personal content out of orchestration inputs and results where references will suffice, because durable execution also persists history. Microsoft documents the execution model in its orchestration guidance.

Stop old workers recreating deleted content

Suppose an embedding worker reads revision 12. An erasure request then blocks revision 12 and removes its chunks. The worker finishes its expensive model call and uploads a new vector from the old text. Passing a cancellation token helps when you can reach the running work, but cancellation alone cannot establish the deletion invariant. A process can miss the notification, resume later or receive a queued retry. Put the final publication decision behind a controlled writer. Every ingestion job carries the source revision it read. The writer rejects work for blocked or superseded revisions. No other ingestion path should have credentials allowing it to bypass that writer.

A check immediately before an external upsert still leaves a race. Deletion can happen after the check and before the write. EF Core concurrency tokens protect updates to the relational record, but they do not automatically make a vector store request part of the same atomic operation. See Microsoft's concurrency documentation for the scope of that protection. The writer therefore needs a real ordering mechanism. One approach serialises publication and deletion commands per source document. It drains any outstanding provider writes before performing the deletion and records uncertain outcomes for reconciliation. A single writer partition or durable command stream can provide this ordering, provided worker failover and outstanding requests are handled rather than merely assumed away.

Another approach writes new artefacts into a staging state and makes them eligible only after an authoritative revision check. If erasure wins the race, the staged record never becomes available and cleanup removes it. This prevents serving stale content, although the staged copy still exists until cleanup completes. A database lease by itself is insufficient if an expired worker can still write to a provider that does not enforce its fencing token. In that case, you need staging, a controlled writer, or an equivalent mechanism that makes obsolete writes harmless to retrieval and discoverable to cleanup. Before verification completes, settle outstanding writes, reject old work and inspect pending artefacts. Repeatedly deleting records without controlling their writers can continue indefinitely.

Delete exact records from the vector store

For Azure AI Search, use the chunk keys recorded in the manifest. The index might contain several versions and generations of the same source, so scope the operation to every obsolete generation covered by the request. The following adapter fragment uses Azure.Search.Documents. It assumes the index key field is named chunkKey and accepts a bounded batch prepared from the manifest:

using Azure.Search.Documents;
using Azure.Search.Documents.Models;

public sealed class SearchChunkDeletion(SearchClient searchClient)
{
    public async Task DeleteBatchAsync(
        IReadOnlyCollection<string> chunkKeys,
        CancellationToken stopToken)
    {
        if (chunkKeys.Count == 0)
            return;

        var batch = IndexDocumentsBatch.Delete("chunkKey", chunkKeys);

        var response = await searchClient.IndexDocumentsAsync(
            batch,
            new IndexDocumentsOptions { ThrowOnAnyError = false },
            stopToken);

        var failures = response.Value.Results
            .Where(result => !result.Succeeded)
            .ToArray();

        if (failures.Length > 0)
        {
            // Persist per-key outcomes in the task ledger before retrying.
            throw new InvalidOperationException(
                $"Search deletion failed for {failures.Length} chunks.");
        }
    }
}

The important behaviour is inspecting each document result. A successful transport response can contain individual failures. A production adapter should persist those outcomes and retry the appropriate keys rather than treating the batch as one undifferentiated success. Azure AI Search documents key-based deletion, idempotent delete behaviour and delayed physical reclamation in its deletion guidance. Logical removal from lookup and search should be verified separately from disk reclamation. Verification should inspect exact known keys and query for remaining records using tenant, source and revision metadata. Exhaust pagination when discovering matching keys. Aggregate document counts are insufficient because unrelated indexing can change the count simultaneously.

An exact lookup returning the expected document not found result can establish logical absence for that key. Authentication failures, an unavailable service or a mistyped index name must remain verification errors. They do not mean the content disappeared. Repeat this for every relevant physical index, including migration copies and older indexes that are still accessible. An index alias pointing to a clean index says nothing about the contents of the previous one.

Remove source copies and dependent outputs

For Blob Storage, cover the extracted text, OCR output, previews and temporary objects as well as the original upload. Versioning and soft deletion require explicit treatment. Microsoft explains that deleting a current blob does not necessarily remove previous versions in its versioning documentation. Your adapter needs to enumerate the covered versions and snapshots, check retention controls and record what it can actually remove. A successful current object delete cannot establish that historical copies have disappeared. Immutable retention can require a separate assessed outcome rather than an endlessly retried delete. Indexer behaviour also needs attention. Azure AI Search blob deletion detection depends on configuration and the indexer observing the relevant state. Do not rely on deleting a source blob first and assume its chunks will follow. Microsoft describes these conditions in its change and deletion detection guidance.

Cached answers need reverse dependencies from source artefacts to cache entries. When that mapping is incomplete, invalidate a broader partition rather than claiming a precise purge you cannot execute. Short expiry reduces persistence but does not demonstrate that the affected response is unavailable now. Saved conversations, summaries and exports need equivalent treatment. An answer can reproduce personal information even after its original passages are gone. Keep source references when generating these outputs and link personal information provided directly by the user as well.

Logging requires particular care. Prevent raw prompts and document bodies entering routine telemetry where possible. Where copies already exist, include the telemetry provider's supported purge process, retention restrictions and acknowledgements in the operation. Removing the application's copy does not remove the exported trace. For queues, prefer reference only messages. Old messages can then be rejected using the source state. If queue payloads already contain personal text, stopping their execution leaves that text stored in the broker or dead letter queue. Use supported removal, purge or assessed expiry procedures and record their actual outcome.

Propagate the request to processors and recipients

The boundary of your application is not the boundary of the request. OCR services, model providers, observability vendors and exported integrations may have received the information. The GDPR includes obligations to communicate erasure to recipients, subject to specified exceptions, and additional provisions where information was made public. The Irish DPC guidance explains this distinction. Track those steps rather than presenting internal deletion as the entire outcome. Record provider side artefacts when they are created: uploaded files, persistent conversations, stored responses or other resources. Give each external integration an erasure capability in your store inventory. Where deletion requires a support request, track the reference and acknowledgement without copying the personal content into the ticket unnecessarily.

Provider retention settings and contracts determine which copies exist and how they can be handled. Obtain evidence through the mechanisms the provider actually offers. A local API response cannot prove that every internal provider replica or backup has been physically overwritten. RAG ingestion also differs from model training. Removing a retrieved document addresses the retrieval corpus. If data was separately used for fine tuning, training or an evaluation service, that is another processing path requiring its own assessment and remediation. Deleting an embedding does not alter model weights.

Treat backup restoration as another ingestion path

Backups can bring erased information back into production. So can restoring an old search export or rebuilding an index from an archived document store. The UK ICO’s erasure guidance discusses backups explicitly. It describes context-dependent handling, including putting retained backup data beyond use until replacement under an established schedule and being clear with the individual about the outcome. This is UK guidance; the appropriate approach for an Irish or other EU controller needs to be assessed for its circumstances. Implement a restore barrier. Restore into an isolated environment, replay the current erasure decisions, remove or restrict affected sources, reconcile derivatives, rebuild eligible indexes and verify the result before enabling users or background jobs.

The erasure ledger must remain current independently of the old snapshot being restored. Restoring the database and its old ledger together would also restore the belief that the deletion never happened. Use opaque internal references and carefully assessed retention for that ledger. A stable identifier or keyed hash may still be personal data. Keep the minimum information needed to prevent resurrection and demonstrate handling, with restricted access and a defined retention decision. This barrier prevents old copies returning through your restore process. It does not justify ignoring backups or automatically promise immediate physical removal from every backup medium.

Verify the closure of the affected data

Propagation is complete only when the coordinator can account for every target in the assessed scope. Start with the affected sources and versions, then follow their dependency edges until there are no additional covered descendants. Do that discovery again after blocking publication and settling outstanding work. The first manifest may have been built while an answer, OCR response or staged chunk was still being produced. Newly discovered targets reopen the operation and receive their own tasks. Verification needs two perspectives. The manifest tells you which known records should be gone. Store reconciliation checks whether actual storage contains covered records missing from the manifest. Use exact metadata filters and controlled inventories for the latter; approximate similarity search cannot establish closure.

A completion decision should establish that obsolete writers cannot publish, all covered known artefacts have verified outcomes, reconciliation found no unexplained copies in the inventoried stores, and external and retention obligations have recorded statuses. Where a backup expiry or provider acknowledgement remains outstanding, expose that separately from verified live system removal. Avoid calling the operation complete because the queue is empty. Messages may have failed, expired or entered a dead letter queue. Completion comes from the task ledger and verification evidence, not transport activity.

Keep audit evidence small. Operation identifiers, scope references, per-store outcomes, timestamps, attempts and exception decisions are usually more useful than retaining the deleted text. Store object references only to the extent needed and protect them appropriately. You can make a strong claim about identified stores and controlled processing paths. You cannot prove that every instance everywhere is gone if uncontrolled exports, unknown stores or uncooperative external systems remain. Surface those limits and unresolved tasks explicitly. Architecture improves the scope of your evidence, it doesnt make an unbounded claim testable.

Test the failures that can resurrect data

The most valuable tests reproduce the gaps between systems. Pause an embedding worker after it reads a source, accept the erasure operation, then resume the worker. The old revision must remain unavailable and any staged output must be cleaned up. Fail the process after an external delete succeeds but before its acknowledgement is saved. Deliver the task again and verify that it converges safely. Fail one key in a search batch and confirm that the operation remains unfinished until the affected key is resolved.

Create a cached answer, a conversation and an evaluation fixture from the same source. Deleting only its vectors must fail the completion checks. Repeat with an older physical index and a previous blob version. Restore a snapshot taken before the request. Keep serving and workers disabled until the current ledger has been replayed and the restored data reconciled. Then attempt a queued retry from before the deletion and confirm that it cannot republish the obsolete revision.

Also test verification failures. A timeout or forbidden response must not become evidence of absence. Check tenant isolation with identical source identifiers in different tenants, and test a multi-person document where the approved action is replacement rather than whole document removal. These tests exercise the system's claim, covered data stops being used, deletion reaches its recorded destinations, and old processing cannot quietly bring it back.

Design ingestion and erasure together

A RAG system becomes much easier to operate when every transformation has an identity, every published artefact has provenance and every store participates in lifecycle handling. The same metadata that explains where an answer came from can explain which answers depend on a document that must be removed. The same source revision that prevents stale indexing can prevent an old embedding job from defeating an erasure request. The same reconciliation process that finds failed writes can find untracked copies.

For the person making the request, the useful outcome is that the covered information is removed or handled according to the assessed decision, its copies are accounted for and the result is communicated accurately. For the engineer, that outcome comes from a durable workflow with explicit scope, controlled writers, verified propagation and a restore process that honours current decisions. Once those pieces are in place, erasure becomes a lifecycle operation your .NET system can execute and demonstrate, including when a worker crashes halfway through it.