Skip to content
AIAC AI ASSURANCE COUNCIL

Auditing AI systems

IIA Standard 14.1 — Gathering Information for Analyses and Evaluation

Standard 14.1 measures sufficiency by reperformance rather than by volume: information is sufficient when it could enable a prudent, informed and competent person to repeat the work and reach the auditor’s conclusion. On an AI system that bar is unforgiving, and the standard closes with a harder requirement still — evidence that cannot be obtained may itself have to be reported as a finding.

§ 1 — Who it binds

Which engagements this governs

Internal audit function

Every internal audit engagement that reaches an AI system, whether the system is the subject of the work or one of the controls relied on within it. Standard 14.1 binds the internal audit function rather than the organisation being audited: it governs what an auditor may conclude from what an auditor holds. It applies identically to a model built in house and to one consumed as a hosted service, where the proof is someone else’s to give.

§ 2 — In practice

Sufficiency is a reperformance test

Most writing on AI audit treats sufficiency as coverage: how many outputs were sampled, how many months of logs were pulled. Standard 14.1 does not define it that way. It requires the auditor to evaluate whether the information is relevant and reliable and whether it is sufficient — three separate judgements — and it fixes sufficiency to a reperformance test: enough to enable a prudent, informed, and competent person to "repeat the engagement work program and reach the same conclusions as the internal auditor". The question is therefore not how much was gathered. It is whether a competent stranger, handed the file and nothing else, would land where you landed.

What makes that hard for an AI system is not the amount of evidence but its reproducibility. A screenshot of a model output records that an output existed; it does not record what produced it. Run the same input against a different model version, a different temperature or a different prompt template and the output moves, so a file naming none of those holds nothing that reperforms. This is why the version, the decoding settings and the inference date are load-bearing rather than housekeeping — and why a supplier’s declared accuracy figure, of the kind Article 15 obliges a provider to state, remains a management assertion until the engagement has reproduced it on its own population.

The requirement with teeth is the last one: if relevant evidence cannot be obtained, internal auditors must determine whether to identify that as a finding. In practice the profession does the opposite. A vendor declines to disclose the training corpus, or the system is hosted and the inference logs are not the client’s to give, and the scope quietly narrows to the controls surrounding the model. That is a planning decision standing in for an audit judgement, and 14.1 does not permit it to be silent — a determination has to be made and recorded whichever way it goes. Where the barrier is contractual rather than technical it is also usually fixable, which makes it a finding a board can act on: the log duty on a deployer under Article 26 reaches only what that deployer controls, so the access that matters is a term settled before signature rather than after an engagement fails.

Reliability is the property most often collapsed into relevance. A model card, a vendor evaluation report and an operational dashboard supplied by the party running the system bear on almost every objective and are independent of none of them. Because 14.1 asks for both properties, the file should show where a single source was corroborated by a second, and should schedule management’s assertions apart from what the engagement tested. The distinction sounds pedantic until a conclusion is challenged, at which point a position resting on the auditee’s own telemetry is very difficult to hold — the point our note on what AI assurance actually means makes about proof you did not produce yourself.

§ 3 — What discharges it

What a defensible evidence file holds

The artefacts an assessor asks to see, and what makes each one sufficient rather than merely present.

  1. 01

    An evidence register keyed to the procedure it supports

    One line per item, naming its source, the date it was obtained, and whether that source is independent of the activity under review. Relevance, reliability and sufficiency are three judgements, and the register is where each was made rather than assumed.

  2. 02

    A reperformance note written by a second person

    Someone who did not do the fieldwork works from the file alone and records whether they arrive at the same result. It is the only direct test of the standard’s own definition, it costs an afternoon, and almost no function performs it.

  3. 03

    The unobtainable-evidence memorandum and its 14.1 determination

    What was requested, from whom, on what date, what was refused, and the reasoned conclusion on whether the absence is reportable. A scope that was quietly trimmed leaves no such document, which is how the decision escapes review.

  4. 04

    A corroboration record where one source carried a conclusion

    The second, independent source consulted, and what it confirmed or contradicted. Where none was available the file should say so, because an uncorroborated single source is a limit on the conclusion rather than a defect in it.

  5. 05

    Management assertions scheduled apart from tested material

    Model cards, vendor reports and self-attested control descriptions belong in their own schedule. Folded into the working file they acquire an evidential weight nobody granted them, and a challenge will find them first.

§ 4 — Worked example

Worked example — forty letters and no model version

An insurer’s internal audit function tests a control over a large language model that drafts claim-decline letters: every draft is reviewed by a claims handler before it goes out. The team pulls forty letters from the quarter, screenshots each alongside the handler’s sign-off, and concludes the control operated effectively. The vendor pushed a model update in week six, which the team knew about and treated as a supplier matter.

Is the evidence sufficient under Standard 14.1?

No, on two grounds, and the sample size is neither of them. Nothing in the file identifies what generated any of the forty drafts — no version, no template revision, no decoding settings, no timestamp — so not one of them can be produced again, and the mid-quarter update means the forty were probably written by two different systems. That alone defeats the reperformance test. The second ground is subtler: a sign-off evidences that a handler pressed approve, which is precisely the behaviour the control is meant to prevent being casual. Proof that the reviewer read the draft, or would have caught a defective one, is absent. And where the vendor will not release the update’s evaluation results, 14.1 calls for a recorded determination on whether that refusal is reportable rather than a reason to trim the engagement.

§ 5 — What a weak answer looks like

Sample size answering the wrong question

Volume offered as an answer to a reproducibility question. The file holds two hundred sampled outputs, every one reviewed and passed, and a memorandum explaining why two hundred was statistically defensible. Nothing in it says which system produced them or how anyone would obtain them again, so the sample size is a precise answer to a question the standard does not ask. Functions that fail here are rarely careless: they have applied a discipline built for records that sit still to a system that does not.

§ 6 — Elsewhere

Where the logs come from

Where another instrument addresses the same obligation. These are correspondences, not comparisons — the Council does not rank one framework against another.

  • EU AI Act

    Article 12 requires a high-risk system to record events automatically over its lifetime. Those logs are frequently the only trace of an inference an engagement can reach, and what they capture decides what can be reperformed later.

A correspondence indicates that two instruments address the same underlying obligation. It is not a mapping endorsed by either body, not a statement that one satisfies the other, and not a judgement about which is more demanding.

§ 7 — When it applies

Which edition this is

  1. 9 January 2024

    The publication date printed on the Global Internal Audit Standards, 2024 edition — the date the current edition was published, not an effective date and not a compliance deadline. The Standards carry no separate effective-date statement; the 2017 framework they supersede is no longer effective.

§ 8 — Where this is assessed

Where this is assessed

Examined in one credential, in the domains named on each card.

CAIA-AUD-FAssurance

Certified AI Audit Fundamentals

Evidence and working papers · 25% of the paper

For internal and IT audit functions that must plan, execute, and report an audit of an AI system, and stand behind the finding.

Audit

§ 9 — provenance

The provision itself

This page sets out what the instrument requires and what discharges it. The official text is the authority — these go straight to it.