Third-Party and Fourth-Party AI Risk: Your Vendor's LLM Is Your Regulatory Problem
You built a compliant consent flow. You wrote a lawful, specific notice. Your users made a genuine affirmative choice.
Then your analytics vendor added an AI insights feature, and now the data those users consented to share with you is being processed by a model hosted by a company you have never assessed, in a jurisdiction you have not evaluated, under terms you have not read. Under the DPDP Act you remain the Data Fiduciary. The accountability did not transfer with the data.
The AI Supply Chain Has More Links Than You Think
Traditional vendor risk assumes a two-party relationship: you and your supplier. AI supply chains routinely run five or six deep, and each link introduces a party with access to, or influence over, your data.
| Layer | Who they are | The risk they introduce |
|---|---|---|
| Base model provider | Trains and publishes the foundation model | Training data provenance; embedded bias; licence terms on outputs; model deprecation |
| Fine-tuning partner | Adapts the model to a domain, sometimes on customer data | Whether your data trained a model serving other customers; retention of tuning datasets |
| Inference host | Runs the model and serves requests | Data residency; prompt and completion logging; retention periods; tenant isolation |
| Orchestration and tooling layer | Vector databases, retrieval frameworks, agent scaffolding | Where embeddings are stored; whether embeddings are treated as personal data |
| Your SaaS vendor | The product you actually bought | Whether AI features are on by default; whether the DPA covers them |
| Your integration | The application calling all of the above | What you send, what you log, what you retain |
Your contract is with layer five. The regulatory obligation attaches to you at layer six. Layers one through four are your fourth-party risk, and in most vendor estates they are entirely undocumented.

Due Diligence Questions Specific to AI Vendors
Your existing security questionnaire covers encryption, access control, certification and incident response. Necessary, insufficient. These are the questions that address AI-specific risk, grouped by what they protect.
Training data provenance
- What data was the base model trained on, and can you evidence the right to use it?
- Do you use customer data - including prompts, completions, uploads and telemetry - to train or improve any model? Is that the default, and can it be disabled contractually rather than through a settings toggle?
- If we fine-tune, who owns the resulting model artefact, and is it isolated to our tenant?
- Are you able to exclude a specific individual's data from future training on request?
Model behaviour and transparency
- Do you publish a model card or equivalent documentation covering intended use, limitations and evaluation results?
- What bias testing has been performed, against what benchmarks, and when was it last repeated?
- How are we notified of model version changes, and can we pin a version?
- What is your deprecation policy and notice period for retiring a model we depend on?
Data handling
- In which countries is data processed, stored and accessible from - including support access?
- Are prompts and completions logged? For how long? Who can read them? Are they used for abuse monitoring by human reviewers?
- How is tenant isolation enforced at the inference layer?
- On termination, what is deleted, on what timeline, and will you certify it in writing?
Supply chain
- List every sub-processor in the AI path, including the base model provider and the inference host.
- How do we receive notice of a new sub-processor, and do we have a right to object?
- What are your own vendors' commitments on training-data use?
Security specific to AI
- What controls exist against prompt injection in your product's AI features?
- Have you performed adversarial testing or red-teaming? Will you share a summary?
- What rate limiting and anomaly detection protect against model extraction?
The single most revealing question in the list is the sub-processor enumeration. A vendor who cannot readily name their base model provider and inference host has not mapped their own supply chain, which tells you what their answers to everything else are worth.

Contractual Clauses That Actually Bite
Diligence findings that do not become contract terms are notes. These are the provisions worth the negotiation.
No secondary training. The vendor shall not use customer data, including prompts, completions, uploads, embeddings and usage telemetry, to train, fine-tune or evaluate any model other than one exclusively serving the customer. Cover embeddings explicitly; vendors frequently treat them as derived data outside the definition of customer data.
Purpose limitation. Processing only for the purposes specified, aligned to your DPDPA purpose registry and the notice you gave.
Sub-processor control. A current published list, advance notice of additions, and a right to object with a termination remedy. Require the AI supply chain specifically, since generic sub-processor lists routinely omit the base model provider.
Data residency. Named permitted countries for processing, storage and access, with support access included. Sectoral localisation requirements flow down explicitly where they apply to you.
Deletion and certification. Deletion on request and at termination within a defined period, extending to backups, logs, fine-tuned artefacts and derived embeddings, with written certification.
Breach notification aligned to your clock. Your obligation to intimate the Data Protection Board runs to 72 hours. A vendor clause giving them 72 hours to tell you leaves you nothing. Negotiate 24 hours or less, and specify the minimum content of the notification.
Audit rights. Either direct audit or, more realistically, a current independent report plus a right to a targeted assessment on reasonable notice following an incident.
Model change notification. Advance notice of model version changes affecting output behaviour, with the ability to pin a version for a defined period. This one is routinely missed and directly affects your validation evidence: a silently upgraded model invalidates the testing your assessment relied on.
Output warranties and indemnity. Where the vendor's model generates content you will publish or rely on, address ownership, third-party infringement, and allocation of liability.
DPDPA-aligned processing terms. Ensure the agreement reflects the Fiduciary-Processor relationship the Act contemplates, including assistance with Data Principal rights requests and with breach intimation.
Continuous Monitoring, Not an Annual Questionnaire
A point-in-time assessment describes a vendor as they were on the day they answered. AI vendors change faster than any other category in your estate - new model versions, new sub-processors, new features enabled by default, revised terms of service.
What to monitor between assessments:
- Terms of service and DPA changes. Subscribe to change notifications; many providers publish version histories. Diff them rather than reading them.
- Sub-processor list changes. Where the vendor publishes a list, monitor it. Where they do not, that is itself a finding.
- New AI features in existing products. Track release notes for every approved SaaS product. This is the most common route by which an assessed vendor becomes an unassessed risk.
- Security incidents and disclosures. Public breach reports, advisories affecting their stack.
- Model deprecations. Announcements that a model you depend on is retiring.
- Certification status. Lapsed ISO 27001, SOC 2 or ISO/IEC 42001 certificates.
Set the reassessment cadence by tier: vendors processing personal data in a Tier 1 decision path get quarterly review; those handling confidential but non-personal data get semi-annual; the rest annually with continuous monitoring on the signals above.
Fourth-Party Risk: The GPU Cloud You Never Heard Of
The genuinely difficult case. Your vendor is contractually solid, based in a jurisdiction you accept, and answers every question well. Their inference runs on a specialised GPU provider in a third country, on hardware shared across their customer base.
You have no contract with that provider. You may have no visibility into their controls. Yet your data is processed there, and if it is personal data of Indian Data Principals, your Section 16 analysis and any sectoral localisation obligation has to account for it.
Practical approaches, in descending order of strength:
Require flow-down. Contractually require your vendor to impose materially equivalent obligations on their sub-processors, and to remain liable for their performance. This is the only mechanism that scales.
Require enumeration. You cannot assess what you cannot name. Make the current list a contractual deliverable, not a courtesy.
Assess concentration, not just individual vendors. Map which base model providers and inference hosts sit behind your vendor portfolio. Most enterprises discover that fifteen vendors resolve to three underlying providers - which is a concentration risk that no individual vendor assessment reveals.
Constrain by architecture where the risk warrants it. For the highest-sensitivity processing, require a deployment model that removes the question: a private endpoint, a dedicated instance, or self-hosting.
Accept and document. Where the risk is genuinely unavoidable and proportionate, record the assessment and the acceptance at the right level. A documented, reasoned acceptance is a defensible position. Silence is not.
A Vendor AI Risk Scorecard
Make the assessment comparable across the portfolio with a weighted score.
| Dimension | Weight | What earns a high score |
|---|---|---|
| Training data commitments | 20% | Contractual no-secondary-training covering embeddings and telemetry, not a settings toggle |
| Supply chain transparency | 20% | Complete named AI sub-processor list, advance notice, right to object |
| Data residency and isolation | 15% | Named regions including support access; enforced tenant isolation |
| Deletion and portability | 10% | Certified deletion covering backups, logs and derived artefacts |
| Security and adversarial testing | 15% | Evidence of red-teaming; injection controls; extraction monitoring |
| Transparency artefacts | 10% | Model cards, evaluation results, version pinning, change notice |
| Incident readiness | 10% | Notification within 24 hours with defined content; demonstrated process |
Score each vendor, plot against the sensitivity of the data they process, and work the top-right quadrant first. The output is a prioritised remediation list rather than a filing cabinet of completed questionnaires.
Conclusion
The regulatory model is unambiguous: accountability sits with the Data Fiduciary, and it does not travel with the data. Every layer of the AI supply chain beneath your direct vendor is a party processing your data under terms you did not negotiate, and the only leverage you have is the contract you hold with the vendor above them.
That makes flow-down obligations, sub-processor enumeration and continuous monitoring the three controls that matter most. Everything else in this article is elaboration on those three.
Actionable recommendations:
- Ask every AI vendor to name their base model provider and inference host. The answer, and how readily it comes, tells you more than the rest of the questionnaire.
- Cover embeddings and telemetry explicitly in no-training clauses. Vendors routinely treat derived data as outside the definition of customer data.
- Negotiate breach notification down to 24 hours. A 72-hour vendor clause consumes your entire regulatory window and leaves nothing for assessment or drafting.
- Require model change notice and version pinning. A silent model upgrade invalidates the validation evidence your assessment depends on.
- Map concentration across the portfolio. Fifteen assessed vendors resolving to three underlying providers is a risk no single vendor assessment will show you.
Extend vendor risk to the AI supply chain. AI-specific questionnaires, automated scoring, sub-processor and terms-of-service change monitoring, and concentration analysis across your whole portfolio. See how Dedups.ai makes fourth-party AI risk visible.