Skip to content
AIAC AI ASSURANCE COUNCIL

Auditing AI systems

IIA Standard 13.4 — Evaluation Criteria

There is usually no document setting out what an AI system is required to achieve, and the common conclusion — that there is therefore nothing to audit against — is the reverse of what Standard 13.4 says. Where management’s criteria are inadequate, the auditor must identify appropriate ones with the board or senior management, and the standard says where they may come from.

§ 1 — Who it binds

When the duty arises

Internal audit function

Every engagement, before fieldwork begins. Standard 13.4 binds the internal audit function and applies whether or not management has set criteria: the duty to assess adequacy runs in all cases, and the duty to identify criteria is triggered by the answer. It bites hardest on AI work because the organisation has usually documented an intention for the system and never a level it can be measured against.

§ 2 — In practice

Inadequate criteria are a duty, not an exit

Standard 13.4 imposes two duties in sequence, and the profession commonly performs neither. First, assess whether management has established adequate criteria for the activity under review — an assessment with a written answer, not a background assumption. Second, and this is the sentence that matters: if the criteria are inadequate, internal auditors must identify appropriate criteria through discussion with the board and/or senior management. Inadequacy is a trigger, not an exit. An engagement that reports "no policy exists" and stops there has done the first and declined the second.

The standard names acceptable sources, and one of them settles the question practitioners find hardest. Criteria may be drawn from authoritative practices (frameworks, standards, guidance, and benchmarks specific to an industry, activity, or profession) — the clean route to using ISO/IEC 42001, the NIST AI Risk Management Framework or the requirements of the EU AI Act as the yardstick for an AI engagement where the organisation has set none of its own. It is worth being exact about what that does and does not do. It does not place the organisation under those instruments, and it is not a conformity assessment. It supplies a stated, external, checkable account of what good looks like, agreed before fieldwork.

The hard part for an AI system is that a criterion has to be a number and a population. For a deterministic control it is binary — the approval exists or it does not — and adequacy is rarely in doubt. For a probabilistic one it is not enough to say the model must be accurate: accurate to what level, measured on whom, over what window, and what result is unsatisfactory. Almost no organisation has decided this, and it is not the engagement’s decision to take alone, which is exactly why 13.4 routes it through the board. Figures a provider is obliged to declare under Article 15 are often the most defensible place to start, because the supplier has already committed to them.

Supplying criteria raises an independence question, and the discussion route is what manages it. An auditor who invents a threshold, tests against it and reports a failure has graded their own homework and will be told so at the closing meeting. Criteria taken to the board and adopted become management’s criteria; the work then tests against something the organisation owns, and they survive to be used again next cycle. Obtaining that adoption is not administrative tidiness — it is what turns the resulting difference into a finding rather than an opinion.

§ 3 — What a weak answer looks like

Criteria found after the fact

Criteria assembled after the fieldwork to fit what was found. The team tests, forms a view, then reaches for a framework that supports it — usually one nobody in the organisation had heard of before the report. It is rarely dishonest and always visible: the yardstick appears in the report rather than in the planning file, no one signed it off, and management’s first move is to dispute the source instead of the facts. Once that argument starts, the condition stops being discussed at all.

§ 4 — What discharges it

What the criteria file must contain

The artefacts an assessor asks to see, and what makes each one sufficient rather than merely present.

  1. 01

    A written adequacy assessment of management’s own criteria

    The first duty 13.4 imposes, and the one leaving no trace when it is skipped. It names what the organisation has set, what that does not cover, and why an aspiration is not something a system can fail.

  2. 02

    A criteria paper naming the source and authority of each criterion

    One row per criterion: what it requires, where it came from, and on whose authority it applies here. A criterion with no stated provenance is the auditor’s view of good practice and will be treated as exactly that.

  3. 03

    Board or senior-management adoption, dated

    Where the auditor supplied what the organisation lacked, the record of the discussion 13.4 requires and the decision that came out of it. Without it the work tests against something only the audit function believes.

  4. 04

    A threshold statement for every probabilistic criterion

    The number, the population it is measured on, the window, and what counts as unsatisfactory. This is the artefact almost always missing, and its absence is why so many AI findings cannot be graded.

  5. 05

    An applicability note for anything drawn from an external framework

    Which provision, which edition, and the basis on which the organisation accepted it as applying to this activity. Naming a standard is not the same as establishing that the organisation is measured by it.

§ 5 — Worked example

Worked example — a policy with no threshold in it

A hospital’s internal audit function is asked to audit an AI triage tool used in the emergency department. The organisation has an AI policy. It says models must be fair, accurate and explainable, that clinical staff remain responsible for decisions, and that suppliers must hold certification. There are no thresholds anywhere in it, and no document states what the tool is expected to achieve.

Is there anything here to audit against?

Not yet — and recording that in writing is the first required step rather than the end of the engagement. Fair, accurate and explainable are objectives; nothing in them can be failed, which is the working test of an inadequate criterion. Having made that assessment, 13.4 obliges the team to identify criteria with the board or senior management. Four sources are available without inventing anything. The supplier’s declared performance figures, to which it is already committed. The escalation rule clinical governance already applies to human triage, which the tool now sits inside. A false-negative ceiling for the acuity category carrying clinical consequence, set by clinicians rather than by audit. And an authoritative practice for the governance limbs. Taken to the board and adopted with a date, those are the criteria — and the work can now produce something management is able to dispute on the facts.

§ 6 — Elsewhere

Where external criteria come from

Where another instrument addresses the same obligation. These are correspondences, not comparisons — the Council does not rank one framework against another.

  • ISO/IEC 42001

    ISO/IEC 42001 clause 9.2 Internal audit places an internal audit duty inside the AI management system itself. Where such a system is in place, the requirements the organisation set for it are already documented and adopted.

  • EU AI Act

    Article 15 obliges a provider to design a high-risk system for appropriate accuracy, robustness and cybersecurity, and to declare the accuracy metrics. A declared figure is a yardstick the supplier has already accepted.

A correspondence indicates that two instruments address the same underlying obligation. It is not a mapping endorsed by either body, not a statement that one satisfies the other, and not a judgement about which is more demanding.

§ 7 — When it applies

Which edition applies

  1. 9 January 2024

    The publication date of the 2024 edition of the Global Internal Audit Standards. It marks which edition is current rather than any deadline for conformance; no effective date is stated in the Standards, and the 2017 framework is no longer effective.

§ 8 — Where this is assessed

Where this is assessed

Examined in one credential, in the domains named on each card.

AIAC-AUDFAssurance

Certified AI Audit Fundamentals

Audit planning and scoping · 20% of the paper

For internal and IT audit functions that must plan, execute, and report an audit of an AI system, and stand behind the finding.

Audit

§ 9 — provenance