Governance
Checking Whether a Vendor Trains on Your Data by Default
A practical method for establishing whether an AI vendor uses your inputs for model training or product improvement by default: what the phrase actually covers, where the binding answer lives in the contract stack, who in your organisation owns the decision, and the carve-outs that quietly reverse a reassuring trust page.
It usually starts with a renewal. A team has been using an AI drafting tool for months on a self-service plan someone put on a corporate card, procurement now wants it on a proper contract, and a question lands in your inbox: "Can you just confirm they don't train on our data?" You open the vendor's trust page, find a line saying customer content is not used to train foundation models, and feel briefly relieved. Then you read the terms of service and find a right to use "service data" to develop and improve the product, plus a separate clause carving aggregated and de-identified information out of the definition of customer data altogether.
That gap between the marketing claim and the contract is where the work sits. The real question is narrower and harder: what is the default for the specific plan, tenancy and interface your people actually use, who is able to change that setting, and what applied during the months before anyone asked. Answering it properly takes an afternoon and produces evidence an auditor can follow.
What "training on your data" actually covers
Vendors and buyers routinely talk past each other because the phrase bundles several distinct processing activities. Separate them before you ask anything, because a vendor can truthfully deny one while doing the others.
- Training or fine-tuning a shared model — your inputs and outputs contribute to weights served to other customers. This is the activity most "we don't train on your data" statements address.
- Fine-tuning a model private to your tenant — often desirable, but it is still training, and it changes retention, deletion and exit obligations.
- Retention for abuse and safety monitoring — prompts held for a fixed window, sometimes reviewable by staff, usually governed by a separate clause and often not switchable off below enterprise tiers.
- Human review — staff or subcontractors reading content for quality, labelling or evaluation. A vendor may not train on your data yet still have humans read it.
- Aggregated, anonymised or de-identified use — the most common escape hatch. If the contract defines this material as outside "customer data", the protective clause you negotiated may not reach it.
- Telemetry and usage metadata — feature use, error traces, sometimes prompt fragments in logs.
Who is responsible for getting the answer
Your organisation is almost always the controller for the personal data your staff put into the tool, and it stays the controller. The vendor is engaged as a processor, and under Article 28(3)(a) of the UK and EU GDPR a processor may process personal data only on your documented instructions. If the vendor decides, for its own purposes, to train its models on your data, it is determining the purposes and means of that processing — and Article 28(10) is explicit that a processor doing so is treated as a controller in respect of it. That is not a technicality. It means the vendor needs its own lawful basis, its own transparency to data subjects, and its own answer on special category data, and it means your instruction has been exceeded. Purpose limitation under Article 5(1)(b) is the companion point: training is a different purpose from delivering the service you bought.
Internally, the DPO advises and challenges but does not own the decision — Article 38(6) exists precisely to keep the DPO out of determining purposes and means. The accountable owner is the business owner of the system, supported by whoever holds the contract. Procurement secures the wording; the system owner confirms the configuration in production; internal audit tests that the two match. Write those names down, because "the DPO signed it off" is not a defensible allocation.
One point of frequent confusion: a vendor training on your inputs is a data protection and contract issue, not something that converts you into a provider under the EU AI Act. Provider status attaches to the party that develops and places the system on the market under its own name. Article 25 sets out when a deployer takes on provider obligations — broadly, putting your own name or trade mark on a high-risk system, substantially modifying it, or changing its intended purpose so that it becomes high-risk. If you fine-tune a model yourself, look hard at that article. If the vendor trains on you, look at your DPA. Where a high-risk system is in scope, Article 26 also makes deployers responsible for input data being relevant and sufficiently representative, to the extent they control it.
Where the binding answer lives
| Source | What it can tell you | What it cannot |
|---|---|---|
| Trust centre or FAQ page | The vendor's current public position; useful for framing questions | Nothing binding — it can change without notice and rarely names your plan tier |
| Standard terms of service | The default licence the vendor takes over your content, including improvement rights | Whether an enterprise order form overrides it |
| Data processing agreement | Documented instructions, sub-processors, retention, deletion, transfer mechanism | Whether the product is configured to match, and whether "anonymised" data is carved out |
| Order form or enterprise addendum | The clauses that actually govern you, including negotiated no-training commitments | What applied during the trial or self-service period beforehand |
| Admin console settings | The live configuration and who can alter it | Whether the setting is contractually guaranteed or a courtesy the vendor can withdraw |
| Sub-processor list | Which model providers sit underneath and where | The terms between the vendor and those providers, unless flowed down |
The checks that find the real default
- Ask, in writing, for the position for our tenant, on our plan, via each interface we use. Application, API and browser extension defaults frequently differ.
- Ask whether the protection is opt-in or opt-out, and whether it applies retrospectively to data already ingested.
- Search the contract stack for "improve", "develop", "analytics", "aggregated", "de-identified", "anonymised" and "derived data". Read each definition, not just the clause.
- Confirm whether the no-training commitment sits in the DPA or only in a policy document the vendor can amend unilaterally.
- Establish the abuse-monitoring retention window, who can read retained content, and whether that is contractual or configurable.
- Ask whether the commitment flows down to sub-processors and underlying model providers, and request evidence.
- Identify who at your organisation can flip the setting, and whether the change is logged and alerted.
- Reconstruct the trial or proof-of-concept period: which terms applied, what was submitted, and whether personal or confidential data went in.
- Confirm how the vendor handles an erasure request that touches material already used in training.
Limits, and when to take advice
Two limits deserve honesty. First, there is no established, reliable technical means of removing the influence of specific training data from a model that has already been trained. If your data went into a shared model during an unmanaged trial, the realistic remedies are deletion of the retained corpus, a contractual commitment to no further use, and a documented decision recording the residual risk. Promises of retrospective removal from model weights should be tested closely.
Second, regulatory thinking here is still developing. The European Data Protection Board has issued an opinion on personal data in AI models, and the ICO has run consultation work on generative AI, but the treatment of model-embedded personal data, anonymity claims and erasure is being worked out case by case. Under the EU AI Act, obligations on general-purpose model providers, including the public summary of training content, depend on templates and codes of practice administered by the AI Office — check the current version rather than a secondary summary. Where a supervisory authority has not settled a point, record your reasoning rather than asserting certainty.
Take specialist advice where the exposure is material: special category or criminal offence data in prompts, confidential client or privileged material, an international transfer without a sound mechanism, a public sector body facing transparency duties, or a discovery that data has been trained on for a sustained period without a lawful basis. The financial stakes are real — up to 4% of total worldwide annual turnover under the UK and EU GDPR, and under the EU AI Act up to 3% for most obligations and up to 7% for prohibited practices. But the more common cost is quieter: a contract you cannot exit cleanly because nobody established, at the outset, what the default was.
- vendor due diligence
- GDPR
- data processing agreement
- AI procurement
- third-party risk
- ISO 42001