Skip to content
AIAC AI ASSURANCE COUNCIL

EU AI Act

Article 15 — Accuracy, robustness and cybersecurity

Article 15 requires an appropriate level of accuracy, robustness and cybersecurity, held consistently across the lifecycle. The provision that decides most assessments is Article 15(3): the levels of accuracy and the relevant accuracy metrics must be declared in the instructions for use. An undeclared metric definition makes the number unusable.

§ 1 — Who it binds

Who it binds, and from when

Provider

Every high-risk AI system. The declaration duty in Article 15(3) runs to deployers, so it also determines what a deployer can rely on when discharging its own Article 26 duties — a deployer given a bare percentage cannot evidence appropriate use of it.

§ 2 — In practice

Why the number is the hard part

Article 15 reads like an engineering requirement and behaves like a disclosure one. The obligation to achieve "an appropriate level" of accuracy, robustness and cybersecurity is deliberately unquantified — appropriate to the intended purpose, judged in context. What is not discretionary is Article 15(3): the levels of accuracy and the relevant accuracy metrics must be declared in the instructions for use. The Act does not tell you how accurate to be. It tells you that you must say how accurate you are, in terms that mean something.

That second half is where assessments are won and lost, because accuracy is not a property of a system — it is a property of a system, a metric, a threshold and a population. Move the decision threshold and precision and recall trade against each other. Change the class balance of the evaluation set and the headline figure moves without the model changing at all. Relabel the ground truth and it moves again. A number published without the definition that produced it cannot be checked, cannot be relied on by a deployer discharging its own duties, and does not discharge 15(3).

The declaration duty runs downstream, which is what makes this a shared problem rather than a provider-only one. A deployer under Article 26 must use the system in accordance with the instructions for use and assign competent human oversight. A deployer handed "94% accurate" with no metric definition cannot calibrate its oversight to anything, and its own compliance inherits the defect. When Article 9 identifies a risk, this is the article that says whether it was ever measured.

Article 15(4) adds a requirement that has become sharply more relevant with systems that keep learning: feedback loops must be mitigated. Where a system’s own outputs re-enter its training data, performance can drift while measured accuracy stays flat, because the test set drifts with it. The evidence here is a control that detects the condition, not a statement that it does not occur. The NIST AI RMF walkthrough sets out the measurement discipline this article assumes but does not supply.

§ 3 — What discharges it

What the evidence must show

The artefacts an assessor asks to see, and what makes each one sufficient rather than merely present.

  1. 01

    Declared accuracy levels and metrics in the instructions for use

    The declaration has to name the metric, not only the score. Precision at what threshold, over what population, against what ground truth. A figure without its definition cannot be checked by a deployer and does not discharge Article 15(3).

  2. 02

    The test methodology and dataset description behind each declared figure

    An assessor reconciles the declared number to the test that produced it. Where the evaluation population differs from the deployment population, that difference is itself a disclosure — the common failure is a figure measured on curated data and declared as though it were operational.

  3. 03

    Robustness testing against the conditions the system will actually meet

    Distribution shift, degraded inputs, adversarial inputs where relevant. Robustness is a claim about behaviour outside the happy path, so evidence drawn only from the happy path does not support it.

  4. 04

    Feedback-loop mitigation records for systems that continue to learn

    Article 15(4) requires mitigation of feedback loops in continually learning systems. The evidence is a control that detects when the system’s own outputs have re-entered its training data, and a record of it having been checked.

  5. 05

    Cybersecurity measures addressing AI-specific attack surfaces

    The four categories Article 15(5) names: data poisoning of the training set, model poisoning of pre-trained components, adversarial examples or model evasion, and confidentiality attacks or model flaws. Surfaces the Act does not enumerate — model extraction, prompt injection — belong here too where the architecture admits them.

§ 4 — Worked example

Worked example

A vendor sells a CV-screening tool to EU employers. Annex III point 4 makes it high-risk. The datasheet states "92% accuracy" and the marketing site repeats it. The figure came from a held-out test set drawn from the same three clients whose historical hiring data trained the model.

Is the declaration compliant with Article 15(3)?

No, on two counts. The metric is undefined — 92% of what, measured how, at what threshold, against whose ground-truth labels? And the population is undisclosed, which matters more here than the number does: a figure measured on three clients’ historical decisions tells a fourth employer very little, and tells it nothing at all about candidates unlike those three clients’ previous applicants. A compliant declaration would name the metric and threshold, describe the evaluation population and how it differs from a typical deployment, and give subgroup figures where the intended purpose makes them material — which for a hiring tool, it plainly does. The vendor does not need a better model to fix this. It needs to say what it already knows.

§ 5 — What a weak answer looks like

What a weak answer looks like

A marketing figure promoted into a compliance artefact. The number in the instructions for use is the same number on the datasheet and the website, because it was written for a buyer rather than for a deployer who has to calibrate oversight against it. The tell is that nobody can say where it came from: the team that produced the evaluation has moved on, the test set was not retained, and the figure has survived two model versions unchanged — which, if it were a measurement, would be remarkable.

§ 6 — Elsewhere

The same requirement elsewhere

Where another instrument addresses the same obligation. These are correspondences, not comparisons — the Council does not rank one framework against another.

  • NIST AI RMF

    The MEASURE function addresses demonstrated performance under conditions of use. It occupies the same evidential ground as Article 15: what was measured, on what population, and whether the figure holds under the conditions the system actually meets.

A correspondence indicates that two instruments address the same underlying obligation. It is not a mapping endorsed by either body, not a statement that one satisfies the other, and not a judgement about which is more demanding.

§ 7 — When it applies

When it applies

  1. 2 December 2027

    Annex III standalone high-risk systems. Moved from 2 August 2026 by the Digital Omnibus — a sixteen-month extension driven by undesignated national authorities and the absence of harmonised standards, not by any relaxation of the Section 2 requirements themselves.

  2. 2 August 2028

    Annex I embedded high-risk systems — medical devices, machinery, vehicles — where AI Act requirements fold into the existing sectoral conformity assessment. Moved from 2 August 2027.

§ 8 — Exposure

Exposure

€15 million or 3% of worldwide annual turnover, whichever is higher

Article 99(4)(a), reached through Article 16(a): a provider must ensure its high-risk system meets the Chapter III Section 2 requirements, and Articles 9 to 15 sit in that section. Neither Article 9 nor Article 15 is itself enumerated in Article 99(4) — the exposure runs through Article 16. Under Article 99(6) an SME pays the lower of the two figures rather than the higher. The 7% ceiling in Article 99(3) applies only to the prohibited practices in Article 5.

§ 9 — Where this is assessed

Where this is assessed

Examined in 2 credentials, in the domains named on each card.

CAIC-FCompliance

Certified AI Compliance Fundamentals — EU AI Act

High-risk obligations and conformity · 25% of the paper

The EU AI Act as enacted: the risk-tier structure, the roles the Act defines, the obligations attaching to each, and the timeline on which they apply.

Compliance
CAIA-AUD-FAssurance

Certified AI Audit Fundamentals

Evidence and working papers · 20% of the paper

For internal and IT audit functions that must plan, execute, and report an audit of an AI system, and stand behind the finding.

Audit

§ 10 — provenance

The provision itself

This page sets out what the instrument requires and what discharges it. The official text is the authority — these go straight to it.