Skip to content
AIAC AI ASSURANCE COUNCIL

Auditing AI systems

IIA Standards 14.2 and 14.3 — Analyses and Potential Engagement Findings; Evaluation of Findings

A finding is not an observation about a model. Standard 14.2 defines one structurally, as a difference between the evaluation criteria and the condition found; 14.3 then calls for root cause where possible, potential effects, and a significance drawn from likelihood and impact. On an AI system the criteria are usually missing and the cause is usually written as the model.

§ 1 — When it applies

Which edition applies

  1. 9 January 2024

    Publication date of the 2024 edition of the Global Internal Audit Standards. It records when the current edition appeared and imposes no deadline; the Standards set out no separate effective date, and the superseded 2017 framework is no longer effective.

§ 2 — In practice

Five parts, and the two that go missing

The structure is not presentation, it is the definition. Under 14.2 a potential finding arises where there is a difference between the evaluation criteria and the condition — so an engagement holding no criteria cannot produce a finding at all, only a description. This is why so much AI reporting reads as commentary: it establishes a condition in detail, attaches an adverse adjective, and never states what the system was supposed to do. Reviewers push back at exactly that seam, because a condition with nothing behind it is the auditor’s preference, and management is entitled to say so out loud.

Standard 14.3 asks for root cause when possible, and that qualifier gets read as an escape hatch on AI work when it is nothing of the kind. Model behaviour is emergent, so there may be no mechanical account of why one output occurred — but that is not the cause an engagement is looking for. The cause is a control that should have caught it and did not: a change gate admitting a version with no subgroup evaluation, an acceptance threshold nobody set, a monitoring arrangement pointed at latency rather than at the control objective. "The model did it" names a mechanism and stops one step short of anything anyone can repair. The same distinction sits under Article 9, where what carries weight is a dated acceptance decision with an author, not a rating.

Significance comes from likelihood and impact, and here an AI engagement holds an advantage it rarely spends. Where every decision in the period is logged, likelihood is not an estimate — it is a count, taken from the population rather than argued from a sample. That converts a graded opinion into a measured one and changes what management is able to dispute. It does depend on the population having been defined during fieldwork, which makes it a working-paper question before it is a reporting one; see evidence sufficiency.

Two findings go unwritten more often than any others. The first is the one 14.3 makes mandatory: where a significant risk is identified, it must be documented and communicated as a finding, whether or not management already has a project open on it. The second is the access finding — material the engagement could not obtain, which is reportable in its own right rather than a limitation tucked into the scope paragraph. Both are dropped for the same reason, that neither is comfortable to raise, and both are what an engagement file gets judged on afterwards.

§ 3 — Who it binds

When a difference becomes a finding

Internal audit function

Every engagement that reaches a conclusion, whichever way the conclusion goes. Standards 14.2 and 14.3 bind the internal audit function during fieldwork rather than at reporting: 14.2 governs how a potential finding forms out of the analysis, and 14.3 how it is evaluated before it becomes one. Where an AI system produced the condition, both apply unchanged.

§ 4 — What discharges it

What a defensible finding sheet carries

The artefacts an assessor asks to see, and what makes each one sufficient rather than merely present.

  1. 01

    A finding sheet carrying criteria and condition as separate fields

    Two fields, both filled, neither inferable from the other. Where the criterion can only be completed with a paraphrase of the condition, the engagement has found something and has not yet established that anything is wrong.

  2. 02

    A cause analysis that terminates at a control

    It names the control that should have prevented or detected the condition and says what it did instead. An analysis whose closing line is a statement about model behaviour has described a mechanism and left management nothing to change.

  3. 03

    An effects statement separating what occurred from what could occur

    Realised effects are countable and attract less argument; potential effects carry the case for acting now. Conflating the two lets management discount the whole by disputing the speculative half.

  4. 04

    A likelihood count taken from the population, not the sample

    Where every decision in the period is recorded, significance rests on a number rather than on an impression of frequency. The artefact is the query and its result, filed with the finding rather than buried in fieldwork.

  5. 05

    The disposition record for potential findings not elevated

    Because a significant risk is reportable, the decisions that matter most are the ones where something was seen and not raised. A short log of what was considered and set aside is the only proof those calls were made rather than avoided.

§ 5 — Worked example

Worked example — a finding everybody agreed with

An audit of a CV-screening model establishes that applications showing an employment gap of more than six months score materially lower, across the whole of last quarter’s applicant pool. The draft reads: "The model demonstrates bias against candidates with career breaks. Recommendation: retrain the model on a more balanced dataset." The head of HR technology has accepted it and opened a retraining project.

What is wrong with it, given that everybody agrees with it?

It is a condition with a remedy bolted on and nothing in between. No criterion: the organisation set no subgroup performance threshold, so nothing has yet been failed — and Standard 13.4 puts the burden back on the auditor to identify criteria with the board rather than to drop the point. No cause: retraining alters what the model does, which is the condition, and leaves untouched whatever allowed a version to reach production with no subgroup evaluation and no threshold to evaluate against. No effect: nobody has said what happened to the applicants concerned, or whether any decision needs revisiting. And no significance: the pool is fully logged, so the number affected is countable and was not counted. Agreement is not reassurance here. The project will close, the condition will move, and the gate that let it through will still be open for the next release.

§ 6 — What a weak answer looks like

An adjective is not a finding

The adjective standing in for the analysis. A report says the model is biased, opaque or insufficiently governed, sets out the condition at length, and moves straight to a recommendation. Nothing states what the system was required to do, nothing reaches a control that failed, and the significance is a colour. It survives the closing meeting because everyone in the room already thinks the situation unsatisfactory — and it cannot survive anyone who does not, which is the reader a finding is written for.

§ 7 — Elsewhere

Where the criteria usually come from

Where another instrument addresses the same obligation. These are correspondences, not comparisons — the Council does not rank one framework against another.

  • EU AI Act

    Article 9 requires a provider to record a judgement that residual risk is acceptable. Where such a decision exists it supplies both a criterion an engagement can test against and an owner a cause analysis can reach.

A correspondence indicates that two instruments address the same underlying obligation. It is not a mapping endorsed by either body, not a statement that one satisfies the other, and not a judgement about which is more demanding.

§ 8 — Where this is assessed

Where this is assessed

Examined in one credential, in the domains named on each card.

AIAC-AUDFAssurance

Certified AI Audit Fundamentals

Reporting the finding · 15% of the paper

For internal and IT audit functions that must plan, execute, and report an audit of an AI system, and stand behind the finding.

Audit

§ 9 — provenance

The provision itself

This page sets out what the instrument requires and what discharges it. The official text is the authority — these go straight to it.