Skip to main content

Governance

The Due Diligence Questions AI Vendors Struggle to Answer

A practitioner's guide to AI vendor due diligence: the questions that reliably expose gaps in a supplier's evidence pack, who actually holds each duty under the EU AI Act and data protection law, and what a usable answer has to contain before you sign.

The Due Diligence Questions AI Vendors Struggle to Answer

It usually arrives as a link. A business team has run a three-month pilot of an AI tool, the results look good, and someone wants it live for the new quarter. What lands in your inbox is a URL to the vendor's trust centre, a SOC 2 report, an ISO/IEC 42001 certificate, and a data processing addendum they would like signed unchanged. The implicit question is whether you can clear it by Friday.

The problem is rarely that the pack is thin. It is often substantial. The problem is that very little of it answers the questions that determine your exposure, because those questions are about how the system behaves on your data, inside your process, under your instructions — and a vendor's evidence pack is deliberately built to be reusable across every customer. Below are the questions that consistently expose that gap, why suppliers struggle with them, and what an answer has to contain before you can rely on it.

Who holds which duty — settle this before you write a single question

The most common failure in AI due diligence is asking a vendor to evidence an obligation that is yours, or accepting a vendor's assurance on one that was never theirs.

Under the EU AI Act, a supplier that develops an AI system and places it on the market under its own name or trade mark is the provider. An organisation using that system under its own authority in the course of its activity is the deployer. Deployer duties are real and cannot be contracted away: use the system in accordance with the instructions for use, assign human oversight to people with the competence, training and authority to exercise it, ensure input data is relevant and sufficiently representative for the intended purpose to the extent you control that data, monitor operation and suspend use and inform the provider where you identify a risk, and retain the automatically generated logs under your control — for high-risk systems, for at least six months unless other law requires otherwise. Employers must inform workers' representatives and affected workers before putting a high-risk system into use at the workplace, and a narrow set of deployers, including bodies governed by public law, private entities providing public services, and deployers of certain creditworthiness and life or health insurance systems, must carry out a fundamental rights impact assessment.

There is a trap here worth naming explicitly. If you put your own name or trade mark on a high-risk system already on the market, make a substantial modification to it, or modify the intended purpose of a system so that it becomes high-risk, you become the provider of that system and inherit the provider obligation set in full. White-labelling and aggressive fine-tuning are the routes organisations walk into this without noticing.

On the data protection side you are almost always the controller and the vendor the processor for delivering the service. But where a vendor uses your data to improve or train its own models, it is determining the purposes of that processing and acts as a controller for it. That is why "we may use customer data to improve our services" is a substantive finding, not boilerplate.

DutyWho holds itCommon mistake
Conformity assessment, technical documentation, CE marking, EU database registration for a high-risk systemProviderDeployer assumes it must, or can, certify the vendor's product itself
Instructions for use, stated accuracy metrics, known limitations and foreseeable misuseProviderMarketing benchmarks accepted in place of documented performance characteristics
Human oversight in live operation: named, competent people with authority to interveneDeployerTreated as a product feature rather than a staffed control
Relevance and representativeness of input data you controlDeployerDelegated to the vendor by silence
Lawful basis, DPIA, data subject rights, international transfer mechanismControllerAssumed to be covered by the vendor's DPA
Serious incident reporting to the market surveillance authorityProvider, with a deployer duty to inform the providerNo contractual trigger tells you when the vendor would tell you
Operating and internally auditing the AI management systemThe vendor's own organisation; the certificate is issued by a certification bodyCertificate read as assurance about a model

The questions that expose the gap

Which model are we buying, and can it change without telling us?

Many suppliers wrap a third-party foundation model. Ask which model and version serves your requests, whether requests are routed dynamically between models or hosting regions, whether you can pin a version, what notice you get before a version is deprecated, and what regression testing runs before a change reaches production. Vendors struggle because the honest answer is often that routing is an optimisation decision made without customer notice. If the system is high-risk, ask whether the instructions for use and stated performance are reissued when the underlying model changes — the documentation must describe the system you are actually running.

Does our data train anything, by default?

Separate the questions the vendor will conflate: prompts and inputs, outputs, telemetry and metadata, support tickets, and any fine-tuned artefacts. For each, ask whether it is used for model training or improvement, whether that is opt-out or opt-in, whether it is set that way for your tenant today, and whether any human reviewer can read content. Get it in the contract, not the FAQ. A trust-centre page is a claim; a contractual restriction on purpose is an instruction under Article 28.

What is the evidence that it works on data like ours?

Public benchmark scores tell you nothing about your document formats, your customer language mix, or your edge cases. Ask for the metrics the provider actually documents, the population they were measured on, the defined thresholds for acceptable performance, and the known limitations. Then run your own evaluation on a held-back, representative sample before go-live, with error analysis broken down by the groups your process affects. Where harmonised standards are still being developed, a vendor cannot claim presumption of conformity through them, and should not imply otherwise.

What happens when it goes wrong, and when we leave?

Ask what the vendor classifies as an incident, how quickly you are notified, and how that meshes with your own breach clock — a processor must notify the controller without undue delay, and a vague "promptly" in the DPA is not sufficient. On exit, "we delete your data" is incomplete. Ask about derived artefacts: embeddings and vector indices, caches, prompt and output logs, evaluation datasets, and any fine-tuned weights built from your material. Ask for the format and mechanism of return, the deletion timetable, and written confirmation. This is the question that most often produces silence.

Certificates: what they cover and where they stop

An ISO/IEC 42001 certificate says a certification body assessed the vendor's AI management system — its scope statement, Statement of Applicability, internal audit programme, management review, and handling of nonconformity and corrective action — and found it conforming. Read the scope statement first: it may cover one business unit or one product line, and not the service you are buying. Ask which body issued it, whether the certification was carried out under a national accreditation body's accreditation, and when the next surveillance audit falls. Accredited certification against 42001 is comparatively new and practice between bodies still varies, so treat the certificate as evidence of governance discipline, never as a technical assurance about a model's accuracy, fairness or safety.

A SOC 2 report is an attestation over a defined scope and period against selected trust services criteria. Check the period, whether it is Type I or Type II, the exceptions noted in the testing section, the complementary user entity controls it assumes you operate, and how subservice organisations are treated. It, too, is silent on model behaviour and training data.

The limits of this, and when to take advice

Several things here are genuinely unsettled. Parts of the EU AI Act depend on implementing acts, further Commission and AI Office guidance, and harmonised standards that are still in development; the phasing of obligations has also been the subject of legislative amendment proposals, so confirm the current position before relying on a date. Certification body procedures for AI management systems differ, and a scope statement means what that body says it means. None of that is a reason to wait — the deployer obligations and the controller obligations exist regardless of who your supplier is.

Take specialist advice where the answer changes your position materially: if you may be flipping into the provider role, if the use case sits near a prohibited practice, where enforcement exposure could reach up to 7% of global annual turnover for prohibited practices, up to 3% for most other AI Act obligations and up to 4% under UK or EU data protection law; if the deployment involves solely automated decisions with legal or similarly significant effects; or if you operate in a regulated sector with its own supervisory expectations. Contractual terms, in particular indemnities and audit rights, should be reviewed by counsel who has seen how those clauses behave when a system fails.

  • AI Act
  • vendor due diligence
  • third-party risk
  • GDPR
  • ISO 42001
  • procurement

More guides

Start Free AI Compliance Review