Skip to main content

Regulation

Declaring Accuracy and Robustness Metrics Under Article 15

Article 15 of the EU AI Act requires high-risk systems to reach an appropriate level of accuracy, robustness and cybersecurity and to declare their accuracy levels and metrics in the instructions for use. This guide sets out who owes that duty, what a defensible declaration contains, how a deployer can inherit it by rebadging or modifying a system, and what to check before you rely on a supplier's headline accuracy figure.

Declaring Accuracy and Robustness Metrics Under Article 15

A procurement pack lands on the desk. The supplier's datasheet says the model is "94% accurate". Nobody in the room can say 94% of what: which population, which decision threshold, which base rate, measured on data from when. Six months later the system is scoring your applicant pool, the outcomes look nothing like the datasheet, and a rejected candidate has asked for an explanation. The question at that point is not technical. It is: who declared that number, on what evidence, and does the system still perform the way its documentation says it does?

That is the situation Article 15 of the EU AI Act is built around. It requires high-risk AI systems to achieve an appropriate level of accuracy, robustness and cybersecurity, and to perform consistently in those respects throughout their lifecycle. It also requires the levels of accuracy, together with the relevant accuracy metrics, to be declared in the instructions for use accompanying the system. The declaration is the part most organisations under-prepare, because it converts an internal evaluation into a statement that regulators, customers and claimants can hold you to.

Who this duty falls on

Article 15 sits among the requirements for high-risk AI systems, and those requirements are owed by the provider — the party that develops a high-risk system, or has one developed, and places it on the market or puts it into service under its own name or trade mark. The provider designs to the standard, tests against it, records the metrics in the technical documentation, and declares the levels and metrics in the instructions for use. A deployer does not write that declaration.

The trap is that a deployer can stop being a deployer. Under Article 25 you are treated as the provider of a high-risk system already on the market if you put your own name or trade mark on it, if you make a substantial modification to it while it remains high-risk, or if you change the intended purpose of a system — including a general-purpose AI system — so that it becomes high-risk. If any of those apply, the Article 15 declaration becomes yours to make, on evidence you can produce. The original provider then ceases to be the provider of that specific system, but must cooperate, supply the necessary information and give the technical access you reasonably need — unless it has clearly specified that its system is not to be changed into a high-risk one. In practice, "we only fine-tuned it and put our logo on it" is the commonest way an organisation acquires obligations it never budgeted for.

A deployer that remains a deployer still has work. Article 26 requires use in accordance with the instructions for use, which means the declared accuracy conditions define the envelope you are permitted to operate in. It also requires human oversight by people with the necessary competence and authority, input data that is relevant and sufficiently representative to the extent you control it, and notification of the provider and the relevant market surveillance authority where you have reason to consider the system presents a risk or you identify a serious incident.

What the requirement actually asks for

Accuracy, and the metrics that define it

"Appropriate" is deliberately relative: appropriate to the intended purpose and to the state of the art. There is no statutory accuracy threshold in the Act, and anyone quoting you a required percentage is inventing it. What the Act insists on is that the level and the metric are stated together and are traceable to real testing. A figure without its metric, its evaluation population, its operating threshold and its date is not a declaration; it is marketing.

Two artefacts have to line up. The technical documentation under Article 11 and Annex IV is expected to record the validation and testing procedures, the validation and test data and its main characteristics, the metrics used to measure accuracy, robustness and compliance with the other high-risk requirements, information on potentially discriminatory impacts, and dated test logs and reports. The instructions for use under Article 13 then carry the outward-facing statement: the levels of accuracy, robustness and cybersecurity against which the system has been tested and validated and which can be expected, plus any known or foreseeable circumstances that may affect those expected levels. That last clause does heavy lifting — population shift, degraded input quality, cohorts thinly represented in the evaluation set, languages or document formats outside the tested range.

Robustness

Robustness is resilience rather than raw performance. The Act frames it around errors, faults and inconsistencies that may occur within the system or the environment it operates in, particularly through interaction with people or other systems, and points to technical redundancy solutions such as backup or fail-safe plans. It singles out systems that continue to learn after being placed on the market: these must be developed so that the risk of biased outputs influencing future inputs — feedback loops — is addressed with appropriate mitigation. If your vendor cannot describe what the system does when an upstream data feed fails, the robustness limb is unevidenced.

Cybersecurity

The cybersecurity limb concerns resilience against attempts by unauthorised third parties to alter the system's use, outputs or performance by exploiting its vulnerabilities. The Act names the model-layer attack classes explicitly: data poisoning, model poisoning, adversarial examples or model evasion, confidentiality attacks and model flaws. This is not the same ground as your information security certification. An ISO/IEC 27001 certificate tells you about the management system around the model; it does not evidence that the model resists evasion inputs.

What has to be written down, and where

ElementWhere it belongsWho produces it
Accuracy levels and the named metricsInstructions for use, and the technical documentationProvider
Validation and test data characteristics, test logs and reportsTechnical documentation (Annex IV)Provider
Known or foreseeable circumstances affecting expected performanceInstructions for useProvider
Robustness and fail-safe measures; feedback-loop mitigationTechnical documentation, summarised in instructionsProvider
Model-layer cybersecurity measuresTechnical documentation, summarised in instructionsProvider
Observed in-service performance against the declared levelsPost-market monitoring evidence and deployer recordsProvider, with deployer input

Reading a supplier's accuracy claim before you rely on it

  • Name the metric. Accuracy, precision, recall, F-score, AUC and calibration error answer different questions. Ask which is declared and why it is the appropriate one for this intended purpose.
  • Ask for the denominator. What population was the figure measured on, how large was it, and how does it compare with the population you will run the system against?
  • Ask for the threshold. A classifier's headline number moves with its operating point. If you can change the threshold in configuration, ask which threshold the declaration assumes.
  • Ask for disaggregation. Aggregate accuracy hides subgroup failure, and Annex IV expects information on potentially discriminatory impacts.
  • Ask for the date and the version. A metric attached to a model version you are not running is not evidence about your system.
  • Ask what degrades it. If the instructions list no foreseeable circumstances affecting performance, that is a documentation gap, not a sign of a robust system.
  • Ask what happens on failure. Fail-safe behaviour, fallback to human decision, and the escalation route.
  • Check the modification question. Confirm in writing whether anything you plan — fine-tuning, rebadging, a new use case — would make you the provider.

Where the detail is genuinely unsettled

Be honest with your board about three things. First, the harmonised standards that will give "appropriate accuracy" concrete meaning are being developed through European standardisation work and, at the time of writing, the relevant technical standards have not all been finalised and cited in the Official Journal. Conformity assessment against them is therefore a moving target. Second, the Act tasks the Commission with encouraging the development of benchmarks and measurement methodologies in cooperation with metrology and benchmarking bodies; until that matures there is no official benchmark suite to point at. Third, the application dates for high-risk obligations have been the subject of active legislative amendment, and you should verify the current date for your classification from a primary source rather than from any secondary summary, including this one.

What is stable is the exposure. Infringement of provider or deployer obligations can attract administrative fines of up to 3% of total worldwide annual turnover; the 7% tier applies to prohibited practices, not to Article 15. Where the system processes personal data, an inaccurate or unfair automated decision can separately engage UK or EU GDPR, with fines of up to 4%.

The limits of this guidance

This is a general explanation of a duty, not advice on your systems. Whether a particular system is high-risk, whether a change amounts to a substantial modification, and what accuracy level is appropriate for a given intended purpose are all fact-specific judgements that turn on the classification analysis and the technical detail. Sectoral regimes — financial services, medical devices, employment law, product safety — may impose their own performance and validation duties that sit on top of the AI Act, and the interaction is not always tidy. Because standards, benchmarks and implementation timing are still moving, verify the current position against primary sources before committing to a compliance date or a conformity route. Where a system materially affects a person's employment, credit, education, health or access to essential services, take qualified legal advice on the classification and on the declaration you are about to sign.

  • EU AI Act
  • Article 15
  • High-Risk AI
  • Technical Documentation
  • Procurement

More guides

Start Free AI Compliance Review