How to answer
Do you use customer data to train AI models?
The fastest-growing question on vendor questionnaires, and the one where a careless yes is hardest to walk back.
No. Customer data is not used to train any model, ours or a third party's. Where customer content is sent to an AI provider to produce a result, it is sent under terms that prohibit training on it, and [the provider] is named in our sub-processor list.
Square brackets are yours to fill in. FillTrust grades an answer like this answered “no” from your documents when your documents support it.
- Answered from
- Your Data Processing Agreement, your sub-processor list, and your terms of service.
- Evidence to attach
- Increasingly, the relevant clause of your DPA, and the terms your own AI provider gives you.
- Where it is asked
- CAIQ v4.0 · DSP-12SIG LiteBespoke vendor questionnaires
Three years ago this question appeared on almost no questionnaires. It now appears on most, it is often the first thing a security reviewer looks for, and it is asked in several forms: do you train on our data, does any sub-processor, and can we opt out.
Answer for the data, not for your infrastructure. A great many companies answer "no, we do not train models" and mean that they do not run training jobs. If customer content is sent to a model provider whose default terms permit training on API inputs, the true answer to the question as asked is yes. What matters to the reviewer is where their customers' data can end up, and the fact that a third party did the training is not a mitigation.
Name the provider. If customer content reaches an AI provider, that provider is a sub-processor and belongs on your sub-processor list with everyone else. Most commercial API terms today do not train on inputs by default, which makes this an easy and strong answer: name the provider, state the contractual position, point at the list.
Separate training from retention. These are asked together and answered together far too often. Training is whether the data improves a model. Retention is how long it sits on somebody's disk. A provider that does not train but retains inputs for 30 days for abuse monitoring is normal, and stating both is a better answer than a bare no that invites the follow-up.
Say whether there is a choice. Some questionnaires ask whether the customer can opt out. If AI processing is intrinsic to your product, the answer is that it cannot be disabled without disabling the feature, which is a legitimate answer. If it is optional, say where the switch is.
Where this grades. Usually as a confirmed no, which is a finished answer rather than a gap: your DPA states it and the sub-processor list backs it up. If your DPA predates your use of an AI provider, it will not say anything at all, and that is the document to update. This is one of the questions where being unable to answer from a document is genuinely worse than an inconvenient answer, because it is the question a buyer's legal team escalates.
How this one goes wrong
Specific to this question, not general advice.
- Answering no for your own models while a sub-processor trains on the data by default. The question is about the data, not about who holds the GPUs.
- Leaving out the AI provider from your sub-processor list. If customer content reaches it, it is a sub-processor, and omitting it is the finding.
- Confusing training with retention. "We do not train on it" and "they delete it after 30 days" are two different commitments and reviewers now ask for both.