10 min readUpdated

Shadow AI in the Enterprise: The Governance Blind Spot Keeping CISOs Up at Night

Marketing is using an AI copywriter. Finance is running a forecasting model in a notebook on someone's laptop. HR is screening CVs through a third-party tool bought on a departmental card. Nobody told information security about any of it.

This is shadow AI, and it is a harder problem than shadow IT ever was - because the data leaves in the request body, the tool often costs nothing, and the business value is immediate and visible while the risk is deferred and invisible.

Three Categories, Three Different Problems

Lumping all ungoverned AI together produces a policy that addresses none of it well. Separate the categories, because the discovery method and the control differ for each.

Shadow AI is a tool adopted without any governance process: a public chatbot, a free transcription service, an API key issued on a personal account. Characteristics: no contract, no data processing agreement, no logging, and frequently terms of service that permit the provider to use submitted data for model improvement.

Sanctioned AI is a tool that went through review and was approved. The residual risk is drift - approved for one purpose, gradually used for another, with no reassessment.

Embedded AI is a feature added to software you already bought and approved. Your CRM, ticketing system, collaboration suite and video conferencing tool have all added summarisation, drafting or analysis features, typically enabled by default. This is the largest category by data volume and the least visible, because no procurement event occurred. The vendor was already approved; the capability is new.

Most programmes focus on the first category, where the emotional salience is highest. The third category is usually where the most sensitive data is actually flowing.

A person working across multiple laptops, representing discovery of ungoverned AI usage
A person working across multiple laptops, representing discovery of ungoverned AI usage

Discovery: Five Techniques, None Sufficient Alone

No single signal finds shadow AI. Run all five and reconcile.

1. SaaS spend analysis. Pull expense reports, corporate card statements and procurement records, and match against a maintained list of AI vendors. Finds paid tools, including small departmental subscriptions below the review threshold. Misses everything free, which is a large share.

2. Identity provider application logs. Your IdP records OAuth grants and SAML assertions. An employee authorising an AI service against the corporate identity leaves a record here. This is the highest-yield single source in most environments, and it is frequently unexamined because it sits with the identity team rather than with security.

3. Network egress and DNS telemetry. Resolve destination domains against known AI provider infrastructure. Catches direct API use and browser-based tools. Degraded by personal devices and by traffic that never traverses corporate network paths.

4. Endpoint DLP and browser telemetry. Detects large clipboard operations and paste events into known AI domains. This is the only technique that sees what data went, rather than merely that a connection occurred. Deploy it where policy and works-council or employment-law constraints permit, and be explicit with staff that it exists.

5. Vendor feature audits. For every approved SaaS product, review the release notes and the current data processing addendum for AI features added since the original assessment, and determine whether they are enabled by default and whether tenant data is used for provider model training. This is manual, tedious, and the highest-value activity in the list.

To these five, add a sixth that is not technical: the amnesty survey. Ask department heads directly what they are using, with an explicit commitment that disclosure now carries no consequence. Framed as an amnesty, response quality is high. Framed as an investigation, you will get silence and drive usage further underground.

The Risk Taxonomy

Board and executive conversations go better when the risk is named precisely rather than described as "data going to AI."

RiskMechanismRealistic consequence
Data exfiltrationConfidential or personal data submitted in a prompt to a provider with no contractBreach obligations under DPDPA Section 8(6); customer contract violation; unassessed sub-processor
Intellectual property contaminationGenerated content of uncertain provenance incorporated into products or marketingOwnership uncertainty over the output; third-party infringement claims
Training on your dataProvider terms permit use of submissions to improve modelsConfidential material influencing outputs to other customers; effectively irreversible
Regulatory non-compliancePersonal data processed with no lawful basis, no notice, and no recordPenalties; inability to answer a Data Principal access or erasure request
Decision biasUngoverned tool used to screen candidates, price, or assess claimsDiscrimination exposure; under the EU AI Act, employment screening is an Annex III high-risk use
Output relianceStaff act on confidently stated but incorrect outputOperational error; professional liability where advice reaches a customer
Availability and lock-inBusiness process quietly becomes dependent on an unmanaged free tierProcess failure when limits, pricing or terms change

The two most commonly underweighted are IP contamination and decision bias. The first because it surfaces years later in a diligence process; the second because the tool that screens CVs rarely looks like a security problem to the department that bought it.

An Acceptable Use Policy That People Will Follow

A policy that says "do not use AI tools" is a policy that gets ignored, and worse, it removes your visibility into usage that continues anyway. The goal is to make the compliant path the easy path.

The structure that works:

Scope. Define AI tool broadly enough to include embedded features, since staff do not perceive a summarise button in their CRM as "using AI."

Data classification rules. The core of the policy, stated as a simple table rather than as prose.

Data classPublic AI serviceApproved enterprise servicePrivate or self-hosted
Public informationPermittedPermittedPermitted
Internal, non-sensitiveNot permittedPermittedPermitted
Confidential or commercially sensitiveNot permittedPermitted with owner approvalPermitted
Personal dataNot permittedPermitted only for a registered purpose with a lawful basisPermitted with controls
Regulated, secret, or customer data under contractual restrictionNot permittedNot permittedPermitted with named approval

Prohibited uses, stated plainly. Decisions affecting employment, credit, insurance or access to services without documented human review. Generating content presented as human-authored where a disclosure obligation exists. Processing another party's data in breach of contract. Circumventing the approved gateway.

The approved catalogue. A short, current list of tools with their permitted data classes. This is the part staff will actually read, and it must be genuinely useful or the policy fails.

A fast route to add a tool. If approval takes six weeks, people will not ask. A tiered intake - same-day for low-risk internal use on non-sensitive data, full assessment for anything touching personal data or decisions - keeps requests coming to you.

Colleagues working through a policy at a whiteboard session
Colleagues working through a policy at a whiteboard session

Enforcement Without Killing Innovation

The failure mode at both extremes is well documented. Prohibit everything and usage goes dark, taking your visibility with it. Permit everything and you have no defensible position when a regulator or customer asks.

What works is a graduated architecture:

Provide a good sanctioned option first. The single most effective shadow-AI control is a capable, fast, enterprise-grade tool that staff prefer to the alternatives. Enforcement against a poor internal option fails; enforcement in favour of a good one barely needs enforcing.

Route through a gateway. An AI gateway or proxy gives you a single point for logging, data-class inspection, redaction, rate limiting and model selection. It converts an unmanageable set of direct integrations into one control point.

Sandbox for experimentation. A segregated environment with synthetic or de-identified data where teams can evaluate models freely without an approval cycle. Most experimentation does not need real data, and saying so removes the main justification for going around you.

Tier access by role and data class. Not everyone needs access to every model with every data class. Bind approval to the use case rather than to the individual.

Handle embedded features deliberately. For each approved SaaS product, make an explicit enable-or-disable decision per AI feature, record it, and re-review at renewal. Default-on is a decision someone else made for you.

Metrics That Show Whether This Is Working

Track four numbers quarterly, and watch the trends rather than the absolutes.

  • Unauthorised AI tools detected. Falling means governance is working. Rising means either adoption is outpacing you or your discovery improved - and you need to know which.
  • Data classes exposed. Count of confirmed instances where confidential or personal data reached an unsanctioned service, by class. This is the number that quantifies actual harm rather than policy violation.
  • Time from detection to remediation. Measures whether the process functions. A long tail here usually indicates unclear ownership rather than technical difficulty.
  • Sanctioned tool adoption. The proportion of AI usage flowing through approved channels. This is the success metric. If it is rising while unauthorised detections fall, the programme is working; if both fall, you have probably just lost visibility.

Conclusion

Shadow AI is not fundamentally a technology problem or a policy problem. It is a problem of the gap between how quickly business teams can adopt capability and how quickly governance can assess it. Closing that gap by slowing the business is not available as an option; closing it by making the governed path faster than the ungoverned one is.

The organisations that manage this well share a pattern: they discovered before they legislated, they provided a good sanctioned option before they enforced, and they treated embedded AI in existing SaaS as seriously as the tools their staff downloaded.

Actionable recommendations:

  • Run all five discovery techniques, plus an amnesty. Each finds a different population. IdP application logs are the highest-yield single source and are usually unexamined.
  • Audit embedded AI features in software you already approved. This is where the most sensitive data flows, and no procurement event will ever alert you to it.
  • Publish a data-class table, not a prohibition. Staff need to know what they may do. A rule they can apply at their desk beats a policy they must interpret.
  • Ship a good sanctioned tool before enforcing. Enforcement against a poor internal option pushes usage underground and costs you visibility.
  • Measure sanctioned adoption alongside unauthorised detections. Both falling together usually means lost visibility, not success.

Discover, classify, and govern every AI touchpoint. Continuous discovery across SaaS spend, identity grants and network telemetry, consequence-based classification, and a single pane covering shadow, sanctioned and embedded AI alike. See how Dedups.ai closes the blind spot.

Ready to get started?

Start securing your cloud infrastructure and optimising costs today.