Certified AI Audit Fundamentals
Evidence and working papers · 25% of the paper
For internal and IT audit functions that must plan, execute, and report an audit of an AI system, and stand behind the finding.
Auditing AI systems
Standard 14.6 asks the engagement file to do one thing: let an informed, prudent internal auditor repeat the work and derive the same results. Its prescribed workpaper format assumes a population you can extract again and a data source that stays where you left it. Neither holds for a model, and closing that gap is what an AI working paper has to add.
§ 1 — In practice
The prescribed format is where this turns concrete. Among the elements 14.6 lists for a basic workpaper are the description of population evaluated, including sample size and method of selection used to analyze data (testing approach) and the source(s) of data covered in the workpaper. Both were drafted for records — a ledger, an extract, a table you can pull again next year and get the same rows back. Name a monitoring platform as your source and a reviewer knows where to go. Name a model as your source and there is nothing stable to return to, because the artefact that generated the material is versioned, hosted, and replaced on a schedule nobody in the audit function controls.
Four fields the standard never names therefore have to be added, and they belong under source rather than under background. The model or artefact identifier, including the version or weights hash actually in service when the procedure ran. The decoding settings — temperature, top-p, seed, or the equivalent for a non-generative system. The prompt or input template, itself versioned. And the date and environment of the inference. None of it is exotic; all of it is discarded by a paper that captures the answer and not the machine, which is the same failure evidence sufficiency turns on, arriving here in the file rather than in the judgement.
A paper is not finished when it is written. Standard 14.6 requires supervisory review and approval by the chief audit executive, and on an AI engagement that review is the cheapest moment to catch a missing version — a reviewer cannot decide whether a conclusion follows from a paper that does not say what was examined. Review notes are consequently evidence in their own right: who queried what, when, and how it was cleared. A note reading "reviewed, agreed" establishes that the step happened and nothing whatever about whether it worked.
Retention has an AI-specific edge that catches functions with otherwise sound records management. A paper is kept for years; the model version it points at may be decommissioned in months, and a hosted endpoint can be withdrawn without the auditee being consulted. A paper citing an artefact it does not itself contain has quietly stopped being repeatable somewhere in the middle of its retention period. Where the audited system is high-risk in the Union, the log duties reaching the deployer may preserve part of the picture — but that is the auditee’s retention policy doing the work, on the auditee’s timetable.
§ 2 — What discharges it
The artefacts an assessor asks to see, and what makes each one sufficient rather than merely present.
01
Every conclusion traceable to the paper supporting it, and every paper to the objective it serves. Indexing is the part that survives staff turnover, and it lets a reviewer test the chain instead of re-reading the fieldwork.
02
Named reviewer, date, the query raised, and its resolution. A note recording only that review occurred proves the step ran and nothing about whether it caught anything, which is the entire reason the step is in the standard.
03
A named act with a date, distinct from supervisory review and not discharged by a status field in a workflow tool. An assessor looks for the person, not the state transition.
04
Artefact identifier, input set, output set, and the comparison method used to decide pass or fail, held together in one place. This is the addition the prescribed format does not name and the repeatability test cannot do without.
05
Not merely the period for the papers, but confirmation that extracts, model versions and endpoints are kept for as long as the papers are. Attachment beats citation for anything that can be decommissioned by somebody else.
§ 3 — Who it binds
Every internal audit engagement, and every part of one that touches an AI system: 14.6 governs the file rather than the subject matter. It binds the internal audit function, and the duty does not shift where fieldwork was performed by a co-sourced provider or an in-house analytics team — their papers become the function’s papers, and the chief audit executive approves them.
§ 4 — Worked example
A bank’s internal audit function completes an engagement over the alert-scoring model inside its transaction monitoring platform. Workpaper W-14 records: population, alerts raised in the quarter; sample, sixty; testing approach, agree the scored risk band to the disposition; result, no exceptions. The source is given as an extract from the platform. Six months later a regulator asks the function to walk through how the conclusion was reached.
Does the file satisfy 14.6, and what breaks first?
"Alerts raised in the quarter" breaks first. That is a description, not a definition: with no extraction query and no row count, nobody can rebuild the set the sixty came out of, and the platform has since been re-indexed. The source breaks second. The extract named is a file on an analyst’s machine rather than an object held in the file, and the scoring model behind it was retrained twice in the intervening months, so the paper cannot say which configuration produced the bands it agreed. What survives is the testing approach, which is stated properly. The remedy is unglamorous and small — the query, the count, the version in service on each date, and the extract attached instead of referenced. None of it would have added a day to the engagement, and all of it would have been there had the review note asked a question.
§ 5 — What a weak answer looks like
The workflow tool mistaken for the file. Fieldwork lives in a GRC platform — a task per step, a status, an attachment or two, a reviewer who clicked approve — and the function reasonably assumes that is the record. What the platform does not capture is most of what the prescribed format asks for: the population as a definition somebody can run again, the source as an object rather than a name, and a review that recorded a question. The work was probably done well. It can no longer be shown to have been.
§ 6 — When it applies
9 January 2024
The date the 2024 edition of the Global Internal Audit Standards was published. It marks the current edition, not a deadline: no compliance date attaches to it, the Standards state no separate effective date, and the 2017 framework they replace is no longer effective.
§ 7 — Elsewhere
Where another instrument addresses the same obligation. These are correspondences, not comparisons — the Council does not rank one framework against another.
EU AI Act
Article 11 and Annex IV require a provider of a high-risk system to hold technical documentation describing how it was built and tested. That is the auditee’s file rather than the auditor’s, and it is often where a reperformance trace can be sourced.
A correspondence indicates that two instruments address the same underlying obligation. It is not a mapping endorsed by either body, not a statement that one satisfies the other, and not a judgement about which is more demanding.
§ 8 — Where this is assessed
Examined in one credential, in the domains named on each card.
Evidence and working papers · 25% of the paper
For internal and IT audit functions that must plan, execute, and report an audit of an AI system, and stand behind the finding.
§ 9 — provenance
This page sets out what the instrument requires and what discharges it. The official text is the authority — these go straight to it.
§ 10 — Also read