11 min readUpdated

Prompt Injection, Model Extraction, and Training-Data Leakage: The New Threat Vector CISOs Must Add to the Risk Register

Your WAF does not catch a jailbreak prompt, because the payload is grammatical English and matches no signature. Your DLP does not flag a model extraction campaign, because each individual query looks like normal use. Your SAST tooling does not find a vulnerability in a model, because there is no code to analyse.

The attack surface changed. The risk register, in most organisations, did not.

The Threat Landscape

Seven attack classes matter for enterprise AI. Each has a distinct mechanism, and grouping them under "AI security" obscures the fact that the controls differ.

Prompt injection. Instructions embedded in content the model processes are interpreted as commands. Direct injection comes from the user typing them. Indirect injection - the more dangerous variant - comes from content the model retrieves: a webpage, an uploaded PDF, an email in a thread, a code comment in a repository. The model has no reliable way to distinguish data it should analyse from instructions it should follow, because both arrive in the same channel as the same kind of token.

The severity scales with the model's capabilities. An injected instruction into a system that only generates text produces bad text. The same injection into a system with tool access - able to send email, query a database, or call an API - produces action.

Data poisoning. Manipulating training or fine-tuning data to alter model behaviour. Includes backdoors, where the model behaves normally except on a specific trigger input. Particularly relevant where training data is scraped, crowd-sourced, or accepted from users through a feedback loop.

Model inversion. Reconstructing characteristics of training data from model outputs or gradients.

Membership inference. Determining whether a specific individual's record was in the training set. This is a privacy attack with direct regulatory consequence - if the training set was a medical cohort, confirming membership discloses a health fact about that person.

Training data extraction. Causing the model to reproduce memorised training data verbatim. Where that data contained personal data, credentials or confidential material, the output is a disclosure.

Model extraction or theft. Systematic querying to build a functionally equivalent model, capturing the value of your training investment without the cost.

Excessive agency. Not an attack in itself, but the amplifier for all of them. A model granted broad tool access, standing credentials, or the ability to act without confirmation converts any successful manipulation into an action with real-world effect.

Code in shallow focus - a prompt injection arrives as well-formed language, matching no signature a scanner holds
Code in shallow focus - a prompt injection arrives as well-formed language, matching no signature a scanner holds

Why Existing Controls Miss These

Understanding the mechanism of the miss is what tells you where to spend.

Perimeter controls inspect the wrong property. A WAF matches patterns associated with injection into structured interpreters - SQL, shell, template engines. A prompt injection is well-formed natural language whose maliciousness is semantic, not syntactic. There is no character sequence to block, because the same words are legitimate in another context.

The trust boundary is inside the payload. Every security architecture separates the control plane from the data plane. Large language models collapse that separation: the system prompt, the retrieved document and the user's question all arrive as one token sequence, and the model's only means of distinguishing them is convention. This is why prompt injection has no complete fix - it is not a bug in an implementation, it is a property of the architecture.

Application security testing assumes deterministic behaviour. A test asserts an expected output for a given input. Model output varies across identical inputs, so a test that passes once does not establish that the behaviour is safe. Testing has to shift from assertion to statistical evaluation over a distribution of adversarial inputs.

DLP inspects content, not intent aggregated over time. Each query in an extraction campaign is individually unremarkable. The signal is in the pattern across thousands of queries - coverage of the input space, systematic variation, sustained rate - which content inspection does not see.

Vulnerability management has nothing to scan. There is no CVE for a model that is too easily manipulated. The relevant knowledge lives in evaluation results and red team findings, not in a package manifest.

Red-Teaming AI Systems

Adversarial testing is the only reliable way to find these before an attacker does. It is a distinct discipline from penetration testing, and outsourcing it to a team that only does network testing produces reassuring, worthless reports.

Scope by capability, not by model. The question is not "can this model be jailbroken" - assume yes. The question is what an attacker achieves once it is. Test the system: the model plus its tools, its retrieval sources, its credentials and its downstream integrations.

Test categories:

  • Direct injection. Instruction override, role-play framing, encoding and obfuscation, multi-turn manipulation building context across a conversation.
  • Indirect injection. Payloads placed in every content source the system ingests - documents, web pages, emails, calendar entries, code repositories, database fields. This is the highest-yield category and the most frequently omitted.
  • Data extraction. Attempts to elicit system prompts, retrieved context belonging to other users, and memorised training data.
  • Tool abuse. Where the system can act, attempt to induce actions outside its intended scope, chained across multiple tools.
  • Bias probing. Systematic variation of protected or proxy attributes in otherwise identical inputs, measuring outcome differences.
  • Robustness. Adversarial perturbations, out-of-distribution inputs, and behaviour at the edges of the input space.

Reference material worth using: the OWASP Top 10 for LLM Applications for the taxonomy, and MITRE ATLAS for adversarial machine learning tactics and techniques mapped in a structure your threat intelligence team will already recognise.

Cadence. On every model version change, on any change to tool access or retrieval sources, and periodically regardless - because the public technique landscape evolves independently of your system.

A person working at a terminal, representing adversarial testing and monitoring
A person working at a terminal, representing adversarial testing and monitoring

Getting These Onto the Risk Register

Register entries expressed in AI vocabulary will not be understood or actioned by the people who own risk. Express them in the structure your register already uses.

Using an ISO 27005 or NIST SP 800-30 style articulation - threat source, threat event, vulnerability, impact:

Threat eventVulnerability exploitedImpactExisting control that partially applies
Adversary embeds instructions in a document the assistant retrievesModel cannot distinguish instructions from dataUnauthorised disclosure; unauthorised action via toolsInput validation (insufficient); output filtering (partial)
Adversary systematically queries the model to reconstruct itNo rate limiting or pattern detection on inferenceLoss of intellectual property; competitive harmAPI rate limiting (usually not tuned for this)
Model reproduces memorised personal data in outputMemorisation of training dataPersonal data breach; Section 8(6) notification obligationData minimisation at training (if applied)
Adversary infers an individual's presence in a training cohortModel overfitting on training dataDisclosure of a sensitive attribute about an identifiable personNone conventionally applied
Compromised training data introduces a triggered backdoorNo provenance verification on training dataIntegrity failure; targeted misbehaviourSupplier assurance (rarely extended to datasets)
Injected instruction causes the model to invoke a tool destructivelyExcessive agency; standing credentialsData loss; unauthorised transactionLeast privilege (rarely applied to model identities)

The right-hand column is the useful one for the register conversation. It shows that these are not entirely uncontrolled - they are partially controlled by existing measures whose scope needs extending, which is a far easier funding case than a new control domain.

Controls That Actually Reduce Risk

Ordered by effectiveness per unit of effort.

1. Constrain agency. The single highest-value control. Give the model the minimum tool access required. Prefer read-only. Require human confirmation for any consequential action. Scope credentials to the invoking user's own permissions rather than to a broad service identity. An injected instruction that can only produce text is an incident; one that can transfer funds is a catastrophe. This control does not prevent injection - it caps the consequence, which is the achievable goal.

2. Isolate untrusted content. Structurally separate retrieved content from instructions using the strongest mechanism your platform offers. Mark provenance explicitly. Treat all retrieved content as hostile by default. Imperfect, and still materially reduces success rates.

3. Filter outputs. Scan completions for personal data, credentials, system prompt content and known sensitive patterns before they reach the user or another system. This is the control that most directly limits extraction and leakage, and it operates independently of how the model was manipulated.

4. Model-access RBAC. Which identities may call which models with which data classes, enforced at a gateway rather than in application code.

5. Rate limiting and anomaly detection tuned for extraction. Not just requests per second. Watch for systematic input-space coverage, unusual query diversity from a single identity, sustained high-volume patterns, and repeated near-identical queries with small variations.

6. Sanitise inputs where the format permits. Strip instruction-like content from retrieved documents, normalise encodings, and remove hidden text - zero-width characters, white-on-white text, HTML comments and document metadata are all standard indirect injection carriers.

7. Canary tokens in training data. Insert unique, meaningless synthetic records into training datasets. If a canary appears in output, you have proven memorisation and extraction. If it appears in a competitor's model, you have evidence of theft. Cheap to implement, and the only reliable detector for several of these attacks.

8. Training data provenance. Know the source of every dataset, verify integrity with hashing, and version rigorously. This is your only defence against poisoning and your only means of scoping the impact if poisoning is discovered.

Monitoring in Production

Three signals, all of which require logging inference in the first place.

Drift detection. Statistical distribution of inputs and outputs against the validated baseline. Distinguish natural drift from adversarial manipulation by correlating with source identity - drift concentrated in traffic from one identity is a different problem from drift across the population.

Anomaly scoring on inference logs. Per-identity behavioural baselines. Alert on deviation in query volume, prompt length distribution, token entropy, tool invocation patterns and rate of refusals triggered. A spike in refusals from one identity is one of the cleanest signals of an active jailbreak attempt.

Feedback loop auditing. Where user feedback influences retraining, that loop is an attack path. Monitor for coordinated feedback patterns, and require human review before feedback data reaches a training set.

Conclusion

These threats are not exotic research concerns. They are the practical consequence of deploying a system whose control plane and data plane are the same channel, and whose behaviour is statistical rather than specified.

The most useful mental adjustment is this: you will not prevent prompt injection. Nobody has. What you control is what an injection achieves, and that is determined almost entirely by how much agency you granted the model. Every other control on the list is worth implementing, and none of them substitutes for that one.

Actionable recommendations:

  • Constrain agency first. Least privilege for model identities, read-only by default, human confirmation for consequential actions. It caps the impact of every attack in this article.
  • Test indirect injection, not just direct. Payloads in retrieved documents, emails and web content are the highest-yield category and the most commonly skipped in AI red team scopes.
  • Add these to the existing risk register using its existing structure. Threat event, vulnerability, impact, partially-applicable control. A separate AI risk register gets read separately, which means rarely.
  • Plant canary tokens in training data. The cheapest reliable detector for memorisation, extraction and model theft, and it must be done before training, not after suspicion.
  • Tune rate limiting for extraction patterns. Requests per second misses a campaign that is patient. Input-space coverage and query diversity per identity do not.

Add AI threat scenarios to your risk register. Pre-built threat entries mapped to mitigating controls, continuous assurance on whether those controls are actually operating, and linkage from red team findings through to remediation. See how Dedups.ai brings AI threats into the risk programme you already run.

Ready to get started?

Start securing your cloud infrastructure and optimising costs today.