# Privacy by Design in AI Product Development: Embedding DPDPA into the ML Pipeline

Privacy cannot be a gate at user acceptance testing. Under the DPDP Act, if the purpose was not specified when consent was obtained, the processing is not lawful - and no amount of model quality changes that. A privacy review that happens after the model is trained is a review that can only discover unfixable problems.

This is how to move the work upstream, into sprint planning, into the pipeline, and into the data model itself.

## What Shifting Left Actually Means Here

"Shift left" is used loosely enough to mean nothing. Concretely, for an ML pipeline, it means four things happen before a model is trained rather than after.

**Purpose is specified before data is collected.** Section 6 requires consent limited to the data necessary for a specified purpose. A dataset assembled first and given a purpose later cannot satisfy that, because the necessity test was never applied. The practical implication is that the question "what are we building and why" must be answered in a form privacy can assess before the first ingestion job runs.

**Lawful basis is established per data category, not per project.** A project may draw on data collected under several bases. Blanket project-level approval hides the field where no valid basis exists.

**Minimisation is a design constraint, not a cleanup task.** The instinct in data science is to gather broadly and let feature selection decide. That inverts the legal test: necessity is assessed at collection, and collecting a field you did not need is not remedied by declining to use it.

**Deletion is designed before it is required.** Retrofitting deletion into a pipeline that assumed permanence is one of the most expensive pieces of remediation in this domain, because it touches the warehouse, the feature store, every derived dataset and every model artefact.

![Privacy decisions taken with the squad at design time, not reviewed downstream at UAT](/assets/blog-images/privacy-by-design-ai-ml-pipeline/dpo-embedded-in-ai-squad.jpg)

## DPIAs Inside Sprint Planning

A DPIA that runs beside the delivery process competes with it and loses. Inside it, it is just another kind of work.

**The screening question.** Every story touching data gets three questions at refinement: does this introduce a new data category, a new purpose, or a new recipient or location? Three no answers close it in seconds, and most stories close in seconds. Any yes raises a linked DPIA task in the same backlog, with the same estimation, the same sprint and the same owner.

The point of this design is that privacy work becomes visible in velocity. When it lives outside the backlog it appears free, gets deprioritised, and then blocks a release. When it is estimated and planned, it competes honestly for capacity and teams plan around it.

**The processing manifest.** Every service commits a manifest declaring the personal data categories it processes, the purposes, the lawful basis, retention class and downstream recipients. It lives in the repository next to the code, is reviewed in the same pull request, and is the single artefact that keeps privacy documentation synchronised with reality.

**Pipeline validation.** A build-time check reads the manifest and validates it: does every declared category have a lawful basis, does every purpose exist in the purpose registry, does a DPIA reference exist where the risk threshold requires one, and has a new category appeared that the approved assessment does not cover. Warn on the first offence, fail on undeclared new categories.

Treat this exactly as you treat a failing security scan: a routine engineering signal with a documented fix, not an external interruption from a compliance function.

## Technical Controls Through the Pipeline

**At ingestion.**

Collect the minimum. Enforce it in code by declaring an allow-list of fields per source and dropping the rest at the boundary rather than downstream - a field that never enters the pipeline cannot leak from it. Tag every record at ingestion with purpose, consent reference, retention class and sensitivity. Reject untagged personal data as a pipeline failure.

**Before training.**

Apply the strongest de-identification the use case tolerates, and understand the differences:

| Technique | What it does | Reversible | Regulatory effect |
|---|---|---|---|
| Pseudonymisation | Replaces identifiers with tokens | Yes, with the mapping | Still personal data; obligations continue |
| Anonymisation | Removes identifiability irreversibly | No | Outside the scope of personal data, if genuinely irreversible |
| Aggregation | Statistics above a group threshold | No | Generally outside scope with adequate thresholds |
| Synthetic generation | Artificial data preserving statistical properties | No, if properly generated | Outside scope, subject to leakage testing |
| Differential privacy | Calibrated noise bounding individual contribution | No | Strong, mathematically stated guarantee |

The common error is treating pseudonymisation as anonymisation. Tokenised data with a mapping you hold is personal data and carries every obligation. Anonymisation is a high bar - if re-identification is possible by combining with other data you or a third party hold, it has not been achieved.

**Filter by consent state at dataset build.** The training set is assembled only from records whose consent is currently valid for a purpose covering model development. This makes consent enforcement a property of the build rather than a manual check, and it is the single most important control in this section.

**In the feature store.**

Every feature carries metadata: source data categories, sensitivity, consent basis, retention class, and whether it is a proxy for a protected attribute. That last field prevents the common failure where a feature such as postcode or institution name imports a protected characteristic that nobody declared.

Governance follows from the metadata: features derived from a purpose cannot be used in a model serving a different purpose without reassessment, and consuming a feature inherits its retention class.

**At training.**

Record lineage: which dataset versions, which consent identifiers in aggregate, which code version, which parameters, which location. Without this, the questions "was this model trained on data we were permitted to use" and "which models are affected by this withdrawal" are unanswerable.

**At inference.**

Log with care. Prompt and completion logs contain whatever the user supplied, which may be anything. Apply retention limits, access control and redaction to inference logs with the same rigour as the source database - and remember that these logs are frequently the least governed personal data store in an AI system.

**After inference.**

Automated retention enforcement across the warehouse, feature store, logs and derived datasets. Deletion that runs on a schedule and is evidenced, not deletion that exists as a policy statement.

![Development work on a monitor, representing minimisation enforced at ingestion](/assets/blog-images/privacy-by-design-ai-ml-pipeline/data-minimisation-at-ingestion.jpg)

## Consent-Aware Architecture

The requirement that withdrawal be as easy as consent has a consequence engineering teams rarely see coming: withdrawal has to reach the model.

The propagation chain:

1. Withdrawal is captured and published as an event.
2. Operational systems suppress the record.
3. The warehouse marks it excluded from future dataset builds.
4. The feature store invalidates derived features for that individual.
5. Any retrieval or personalisation index removes it, so the individual cannot be surfaced at inference time.
6. A retraining trigger is evaluated.

Step six is the unresolved one. Deleting a training record does not remove its influence from an already-trained model, and there is no complete technical remedy available today.

Defensible positions, in order of preference: avoid the problem by training on aggregated, anonymised or synthetic data where the use case permits; maintain dataset-to-consent lineage so the problem can at least be scoped; define a retraining trigger on a threshold of withdrawn records or a fixed interval; suppress the individual at the output and retrieval layers in the meantime; and document the position with DPO review.

The organisations that will handle this acceptably are the ones recording lineage now. Reconstructing which consent identifiers contributed to which dataset version across several past model generations is not realistically achievable after the fact.

## Documentation as Code

Privacy documentation goes stale because it lives somewhere the code does not. Move it into the repository and it stays current for the same reason the code does.

**Model cards** covering intended use, out-of-scope uses, training data description, evaluation results including disaggregated performance, limitations and ethical considerations. Committed alongside the model, versioned with it, rendered from source rather than maintained separately.

**Data cards** per dataset: composition, collection method, consent basis, preprocessing, known biases, retention and permitted uses.

**Lineage graphs** generated from pipeline definitions rather than drawn by hand, showing the path from source through transformation to model to inference.

Generate the compliance view from these artefacts. A DPIA that reads the manifest, the data cards and the lineage graph is a DPIA that cannot drift from the system it describes, because it is derived from it.

## The Collaboration Model

The structural change that makes the rest work: the DPO participates in the squad rather than reviewing its output.

**Embedded, not downstream.** The DPO or a privacy engineer attends refinement and design reviews for teams building AI on personal data. The cost is a few hours a week. The saving is the launch that does not get blocked in its final week over a design decision made in month one.

**Privacy champions.** One engineer per squad with additional training, empowered to raise the screening questions without waiting for the DPO. This scales the function without scaling headcount, and it changes the dynamic from external review to internal ownership.

**Shared definition of done.** Manifest committed, DPIA reference where required, retention implemented, deletion tested. Not a separate privacy checklist - the same definition of done the team already uses, extended.

**A fast path for low-risk work.** If every change requires DPO review, the DPO becomes a bottleneck and teams route around them. Publish clear criteria for what proceeds without review, and mean them.

## Conclusion

Privacy by design in an ML context comes down to a small number of decisions made early: what you collect, how you tag it, what basis you rely on, how you will delete it, and whether you record enough lineage to answer questions later.

Each is cheap at design time and expensive afterwards. The most expensive of all is lineage, because it is the only one that cannot be reconstructed retrospectively - and it is the one that determines whether you can ever answer a withdrawal request about a trained model.

**Actionable recommendations:**

- **Put the processing manifest in the repository and validate it in the pipeline.** It is the single artefact that stops privacy documentation drifting from what the system actually does.
- **Filter training datasets by consent state at build time.** This converts consent enforcement from a manual check into a property of the build, and it is the highest-value control in the pipeline.
- **Record dataset-to-consent and model-to-dataset lineage from today.** It cannot be reconstructed later, and without it the withdrawal question has no answer at all.
- **Tag features as proxies for protected attributes.** Postcode and institution name import protected characteristics that nobody declared, and aggregate accuracy conceals the consequence entirely.
- **Embed the DPO in refinement, and publish a fast path.** Review at the end blocks launches; review at the start costs hours. A bottleneck with no fast path just gets routed around.

> **Put privacy checkpoints inside your CI/CD and MLOps pipeline.** Manifest validation, automated DPDPA control checks at build time, consent-filtered dataset builds, and lineage captured as a by-product of the pipeline rather than as documentation effort. See how [Dedups.ai](https://dedups.ai) shifts privacy left.
