Skip to main content

Operations

Why the Vendor Demo Passes and Your Own Data Fails

A practical guide for compliance, risk and DPO teams on closing the gap between an AI vendor's demonstration and performance on your own records: what the demo was really showing, who holds which duty under the EU AI Act and data protection law, how to design an acceptance test that produces defensible evidence, and where pilots usually go wrong.

Why the Vendor Demo Passes and Your Own Data Fails

The pattern is familiar. A vendor runs a demonstration for the business sponsor, the outputs look accurate and well-judged, and procurement moves to contract. Six weeks into the pilot the operational team reports that the system misreads your document formats, misclassifies your highest-volume case type, and produces confident answers that a reviewer has to unpick. Nobody lied. The demo did what it showed. It simply was not a test of your organisation.

This usually lands on the compliance or risk desk at the worst moment: after commercial expectations are set, and often after someone has already written a business case with a benefits figure attached. You are asked either to bless a deployment on the strength of vendor evidence, or to explain what would be enough. That question has a defensible answer, and it is worth having ready before the demo rather than after it.

What the demonstration was actually showing

A demonstration is a controlled exhibition of capability. It is not a performance claim about your data, and in most cases the vendor has not represented it as one. The gap between the two is created by ordinary, non-sinister differences.

Demo conditionYour production conditionWhat to test
Inputs chosen by the vendor from clean, complete examplesScanned attachments, mixed formats, missing fields, free-text notesA sample drawn by you from live records, including the messy tail
A retrieval corpus curated for the demonstrationYour policies, contracts and case history, including superseded versionsBehaviour when the source material is stale, contradictory or absent
Prompts, thresholds or settings tuned to those examplesDefault configuration, then whatever a business user changesPerformance at the configuration you will actually deploy
A specific model version on the dayA model the vendor may update during the termVersion pinning, change notice, and a retest trigger in the contract
Light load, one language, recent dataPeak volumes, multiple languages, historic recordsThroughput, latency and subgroup performance under real distribution
An expert operator steering the sessionStaff working at pace with limited contextOutcomes with your actual users, not the sponsor

None of this makes the product unsuitable. It makes the demonstration the wrong evidence to rely on for a deployment decision, and that distinction is what you need to record.

Who is responsible for the gap

This is where allocations are most often confused, and getting it wrong in a contract or a risk assessment is expensive.

Under the EU AI Act, the organisation that develops a system and places it on the market or puts it into service under its own name or trade mark is the provider. An organisation using the system under its own authority in the course of its activity is the deployer. For a high-risk system, the provider carries the data governance duties over training, validation and testing data, and must supply instructions for use that state the system's capabilities, its limitations, the declared level of accuracy and the metrics used, and the circumstances that may affect performance. The deployer must use the system in accordance with those instructions, assign human oversight to people with the competence, training and authority to exercise it, and — to the extent it controls the input data — ensure that input data is relevant and sufficiently representative in view of the intended purpose. That last duty is precisely the demo-versus-your-data problem, and it sits with you, not the vendor.

Article 25 matters here too: a deployer that puts its own name or trade mark on a high-risk system, makes a substantial modification, or changes the intended purpose can itself become the provider, with the full provider obligation set. Rebadging a vendor tool as your own internal product is not a cosmetic decision.

DutyWho holds it
Data governance over training, validation and testing sets (high-risk)Provider
Declaring accuracy levels, metrics and known limitations in the instructions for useProvider
Post-market monitoring of the system across its user baseProvider
Ensuring input data is relevant and sufficiently representative for the intended purposeDeployer, where it controls the input
Human oversight in operation, and monitoring use against the instructionsDeployer
Determining the purposes and means of the personal data processingController (normally you)
Processing only on documented instructions from the controllerProcessor (normally the vendor)
Certifying the management system against ISO/IEC 42001Certification body, not the organisation and not the vendor

On the data protection side, if you supply live records for a pilot you are almost certainly the controller for that processing and need a lawful basis, a processing agreement, and — where the processing is likely to result in high risk — a DPIA before you start. If the vendor uses your pilot data to improve its own models, it is pursuing its own purpose and is a controller for that activity, which is a different legal position from the one your contract probably describes. Enforcement exposure is real: most EU AI Act obligations carry fines of up to 3% of global annual turnover, and UK and EU GDPR breaches up to 4%.

Designing an acceptance test that produces evidence

The point of the exercise is not to prove the vendor wrong. It is to generate a record that a regulator, an internal auditor or your own successor can read. Work through this before any data moves.

  • Draw the sample yourself, from your own live or historic records, stratified so that low-frequency but high-consequence cases are represented rather than averaged away.
  • Set the pass threshold in writing before you see any results, expressed in the metric that matters operationally — usually a false-negative or false-positive rate on a specific decision, not overall accuracy.
  • Establish the human baseline. If you do not know how your current process performs on the same sample, you cannot say whether the system is an improvement.
  • Define the error taxonomy up front so that failures are categorised consistently and not renegotiated case by case.
  • Have someone independent score the outputs — not the sponsor, not the vendor's implementation consultant.
  • Check subgroup performance where the decision affects individuals, and record what you found even if the numbers are thin.
  • Test the failure and escalation path, not only the happy path: what a user sees when the system is uncertain, and whether they can realistically override it.
  • Pin the model version, record it, and agree in the contract what notice you get before it changes and what triggers a retest.
  • Confirm in writing whether your inputs and outputs are used for training or fine-tuning, and whether that setting is on by default.
  • Keep the sample, the scores, the threshold and the sign-off as a retained record with a named owner.

If you operate an ISO/IEC 42001 management system, this is where it earns its keep. Your scope statement should make clear whether procured third-party AI systems are inside the boundary; your Statement of Applicability records which Annex A controls you apply and justifies any exclusions; and pre-deployment validation should be a documented requirement that internal audit can sample against. Deploying on vendor evidence alone, where your own procedure requires validation on representative data, is a straightforward nonconformity — one that internal audit should raise, corrective action should address at root cause, and management review should see. ISO/IEC 42001 certifies a management system, not a model; no certificate, yours or the vendor's, tells you how the system performs on your records.

Where this usually goes wrong

The recurring failures are procedural rather than technical. The sample is supplied by the vendor, or drawn from the easy queue because that data was readily available. Results arrive without a pre-agreed threshold, so the discussion becomes an argument about whether the errors "really matter". Testing runs on synthetic or heavily redacted data that no longer resembles the live distribution. Success is measured by user enthusiasm in a workshop rather than by outcomes. A pilot passes, the vendor ships a model update, and nobody retests. Or the pilot quietly becomes production without a decision point, which means no record exists of who accepted the residual risk.

Limits, and when to take advice

Several things here are genuinely unsettled and it is more useful to say so than to pretend otherwise. The harmonised standards that will define presumption of conformity for high-risk systems, and much of the Commission's implementing guidance, are still in development, so "how much testing is enough" has no fixed answer yet; document your reasoning and your threshold, because the reasoning is what you will be asked to defend. Application dates under the EU AI Act are staged and vary by obligation, so confirm the timetable applicable to your system rather than assuming a single deadline. The UK has no equivalent AI statute in force and its regulators apply existing law — for most of these questions, the ICO under UK GDPR and the Data Protection Act 2018 — so verify the current UK position before relying on it. Certification body practice under ISO/IEC 42001 also varies, and your body's own audit procedure governs what evidence it will accept.

Take specialist advice where the system makes or materially influences decisions about individuals with legal or similarly significant effects, where you are considering putting your own brand on a third-party system, where sector regulation applies on top of the general framework, or where a pilot has already processed live personal data without a completed DPIA. Those are situations where a documented judgement made early is far cheaper than a remediation programme later.

  • AI procurement
  • vendor due diligence
  • EU AI Act
  • acceptance testing
  • DPIA
  • ISO 42001

More guides

Start Free AI Compliance Review