AI Vendor Diligence 2026: The Questions Your Security Questionnaire Doesn't Ask
Standard vendor review confirms encryption at rest, an audit report, and an incident process. None of that tells you who owns the output, whether the indemnity reaches your use case, or what happens when the model version you validated is retired.
Why the Existing Questionnaire Misses
Vendor security review evolved to answer one question: if this vendor is breached, what happens to our data. It does that well. The risks that AI vendors introduce are mostly not breach risks — they are ownership risks, provenance risks, and validity risks, and a questionnaire built around confidentiality does not have fields for them.
The gap shows up in a specific way. A team runs the standard review, the vendor passes, the tool is deployed, and eighteen months later someone asks whether the marketing copy the tool produced can be registered, whether the model that scores applicants was ever tested for disparate impact, or whether the vendor's terms permitted them to train on the customer records the integration has been sending since launch. All three are answerable at procurement and expensive afterward.
What follows is the delta — the questions to add, why each one matters, and what a non-answer tells you.
Four Areas the Standard Review Does Not Cover
Output ownership and rights
Ask who owns generated output, whether the vendor asserts any license to it, whether identical output may be delivered to other customers, and what the vendor represents about the protectability of the output.
Why it matters at signature: Purely machine-generated material sits outside the scope of what can be registered for copyright in several major jurisdictions, which affects whether you can enforce rights in what you produce. If output is going into brand assets, product content or anything you intend to protect, this determines whether you have an asset or just a file.
Training-data provenance and infringement exposure
Ask about source categories, respect for scraping opt-out signals, use of any known-infringing dataset, pending or threatened claims, and whether the vendor will warrant sufficient rights in the training corpus.
Why it matters at signature: This is the live litigation frontier. You are unlikely to get a manifest, but the shape of the answer is diagnostic: a vendor that characterizes sources and stands behind a warranty has priced the risk, and a vendor that declines to characterize them is asking you to absorb it silently.
Input reuse and retention
Ask whether inputs and outputs are used for training, fine-tuning, evaluation or abuse monitoring; what the retention period is for each purpose; whether humans review retained content; and whether the restrictions flow down to subprocessors and survive termination.
Why it matters at signature: Abuse-monitoring retention is the most commonly overlooked copy of customer data in the stack. It is frequently excluded from no-training commitments, retained on a separate schedule, and subject to human review — which changes your privacy notice, your DSAR scope and your confidentiality analysis.
Model lifecycle and validity
Ask for the deprecation notice period, availability of prior versions during transition, notice of material behavioral changes to a pinned version, and the vendor's evaluation methodology and results for your use case.
Why it matters at signature: Any validation you perform — accuracy testing, bias auditing, a regulatory assessment — is evidence about a specific model version. If that version can be replaced without notice, your evidence expires without notice too, and you may be operating an unvalidated system in a regulated decision path.
Reading the Indemnity Properly
Copyright indemnification became a standard marketing feature, and the marketing describes the grant while the contract operates through the exclusions. The five that appear most often:
- Customer-supplied input. Output derived from material you provided is excluded. Since retrieval-augmented and document-grounded generation is what most enterprise deployments actually do, this single carve-out can remove the indemnity from your entire use case.
- Filters and safety features. Coverage is conditioned on running the vendor's content filters at default settings. Teams that loosened a filter to reduce false refusals have often forfeited the indemnity without recording the decision anywhere.
- Modified output. Editing generated text before publication can void coverage. Since nobody publishes raw output, verify how this exclusion is scoped — some are narrow, some swallow the grant entirely.
- Prompt-driven claims. Where the prompt requested content in the style of a named creator, or referenced a specific work or brand, the claim is typically excluded and attributed to your instruction.
- Caps and defense control. An indemnity capped at fees paid in the trailing twelve months is not meaningful protection against a statutory-damages claim. Check the cap, and check whether you control the defense or are bound to the vendor's choices.
The test to apply: write down your top three use cases in one sentence each, then trace each one through the exclusions and record whether it is covered. If none of the three survives, the indemnity is a marketing feature and should be treated as one when pricing the risk.
The Buyer-Side AI Diligence Checklist
Add these to the standard security review. The right-hand consequence of a weak answer is a negotiation position, not automatically a rejection.
- ☐Who owns generated output, and does the vendor retain any license to it?
- ☐Can substantially identical output be delivered to other customers, including competitors?
- ☐What does the vendor represent about training-data sources and rights held?
- ☐Are there pending or threatened claims relating to training data or model output?
- ☐Will the vendor warrant sufficient rights, and what survives if that warranty is breached?
- ☐Are inputs or outputs used for training, fine-tuning or evaluation — contractually, not by setting?
- ☐What is retained for abuse monitoring, for how long, and does a human review it?
- ☐Does the restriction flow down to every subprocessor and survive termination?
- ☐What is the deletion mechanism and turnaround for a consumer request routed through you?
- ☐Where is data processed and stored, and which subprocessors are named today?
- ☐What evaluation results exist for a use case resembling yours — not a generic benchmark?
- ☐Has the model been tested for disparate impact where it touches a regulated decision?
- ☐What is the deprecation notice period for a version you have validated?
- ☐Will you be notified of material behavioral changes to a pinned version?
- ☐Can you pin a version, and for how long is a prior version available during transition?
- ☐Trace each of your top three use cases through the indemnity exclusions in writing
- ☐Check the liability cap against a realistic claim, not against annual fees
- ☐Secure audit rights and access to model documentation before signature
- ☐Require notice of any change to the subprocessor list with a right to object
- ☐Define exit: data return format, deletion certification, and transition assistance
Frequently Asked Questions
We have no leverage — the vendor won't negotiate their standard terms. What now?
Then price the risk and record the decision, which is a legitimate outcome. Document which questions went unanswered, which exclusions apply to your use cases, and what you are accepting as a result, with a named owner and a review date. This does three useful things: it converts an invisible exposure into a tracked one, it gives you a trigger to revisit at renewal when leverage is different, and it demonstrates a governance process if anyone later asks how the decision was made. Separately, look at whether the use case can be reshaped to reduce the exposure — moving a tool out of a regulated decision path is often easier than moving a vendor's contract terms.
Is a vendor's SOC 2 report relevant to any of this?
It is relevant to security controls and largely silent on everything in this article. A SOC 2 examines whether the vendor operates the controls it says it operates against the trust services criteria. It does not evaluate training-data provenance, output ownership, model fairness, or deprecation practice, because those are not what the framework covers. Treat it as a necessary input to the security half of the review and as no evidence at all on the AI-specific half. The pattern to avoid is a review that concludes at 'they have a SOC 2' — that answer is about a different set of questions than the ones that will produce your next surprise.
How deep should diligence go for a small tool, like a writing assistant a team wants?
Scale it to where the data goes and what the output does, not to the contract value. A cheap tool that a support team pastes customer conversations into is a higher-exposure decision than an expensive one that never touches personal information. For low-cost tools, a short screen usually suffices: does customer or personal data enter it, is that data used for training, does the output enter a regulated decision or a published asset. Any yes escalates to the full set. The most common failure in this category is that the tool never reaches procurement at all — someone expenses it — so the screen has to be paired with a way of discovering tools already in use.
Should we require the vendor to tell us when they change subprocessors?
Yes, and pair it with a right to object. AI vendors change inference providers, hosting regions and safety-review contractors more often than traditional SaaS vendors change infrastructure, and each change can alter where your data is processed and who can see it. Without notice, your own privacy disclosures and transfer assessments silently go stale, and the first you learn of it is when a customer or regulator asks a question you cannot answer. A notice period with an objection right — even one that ultimately gives you only a termination remedy — keeps your documentation accurate and gives you a decision point.
What single question is most diagnostic if we only get to ask one?
'Are our inputs and outputs used for training, fine-tuning, evaluation or abuse monitoring — and where is that stated in the contract rather than in a settings page?' It is diagnostic on several axes at once. The substance matters directly for confidentiality and privacy. The form of the answer tells you whether the vendor has thought carefully about data handling or is reciting marketing language. And insisting on the contractual location rather than the product setting reveals quickly whether their commitments are durable or subject to change with the next release.
Start With the Tools You Already Bought
Most organizations have more AI vendors than their procurement register shows, because the cheap ones arrived on a corporate card. Pull the expense data, list every AI tool in use, and run the four-question screen: does personal data enter it, is that data used for training, does its output reach a regulated decision, and does its output get published.
The tools that answer yes twice are where the diligence set earns its cost — and renewal is the moment you get to ask the questions you skipped at signature.