Auditing AI systems
NIST AI RMF MEASURE 2.3 · IIA Standard 14.1 — AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s)
Sampling was built for records that sit still. Standard 14.1 tells an auditor to consider whether to test a complete population or a representative sample, and observes that data-analysis software has made whole-population testing practical. For an AI system the whole population is usually available and almost never used, which is how a clean sample and a defective model coexist.
§ 1 — Who it binds
Which controls this reaches
Any test of a control whose operation depends on a model output rather than on a rule — scoring, ranking, classification, extraction, generation — including where the control under test is the human step that follows the output. It binds the internal audit function. The MEASURE subcategories are voluntary guidance and bind nobody; they are used here as a source of criteria and of method.
§ 2 — In practice
Sample or population, and why it is usually the wrong choice
Two passages in 14.1’s implementation guidance point the same way. The auditor is told to consider whether to test a complete data population or a representative sample, with the observation that data-analysis software has made examining an entire population feasible; and where a sample is used, it must be as representative of the entire population as possible. An AI system is the case those two sentences were waiting for. Every inference is recorded, a model can be re-run across a period at the price of compute rather than of weeks, and the population is a query. Sampling remains the default anyway, which makes it a habit rather than a decision.
It matters because a set of model outputs and the control under test are not the same object. Forty outputs test forty outputs. The control is the model together with the gate around it, and on a probabilistic system a clean draw and a defective population are entirely compatible: failures concentrate at the margins of the input distribution, and a random draw under-represents margins by construction. Representativeness therefore has to be argued against the deployment population rather than against the log — the difference between a declared accuracy figure of the kind Article 15 requires and a measurement taken on the traffic in front of you.
Two subcategories of the NIST AI Risk Management Framework explain why a supplier benchmark is not proof about a production system. MEASURE 2.3 asks that performance or assurance criteria be measured and demonstrated for conditions similar to deployment setting(s). MEASURE 2.5 asks that limitations of the generalizability beyond the conditions under which the technology was developed are documented. Together they make the deployment-similarity argument an audit artefact rather than a vendor claim: the question is not whether a model scored well, but whether the conditions it scored well under resemble the ones it now works in. A walkthrough of the framework sets out where MEASURE sits against the rest of it.
A test on a probabilistic system also has a shelf life, shortened by two things. Version is the first: a result establishes something about the configuration in service when it ran, and a retraining ends its authority, so a period holding three versions holds three tests rather than one. Reproducibility is the second: the procedure only repeats if the inference conditions were pinned at the time, which is a question of what the paper recorded rather than of how the test was designed. Both are settled during fieldwork and both are usually discovered at review; see evidence sufficiency.
§ 3 — What discharges it
What the test file must show
The artefacts an assessor asks to see, and what makes each one sufficient rather than merely present.
01
A population definition with a row count and the extraction query
The query, the count it returned, and the date it was run. Without those a later reader cannot rebuild the set, and the difference between a population and a selection from it cannot be checked at all.
02
A justification for sampling where the population was available
Standard 14.1 asks the auditor to consider testing everything. Where a subset was taken instead, the paper should say why — cost, access, or a genuine constraint — rather than leave the default unexamined.
03
A test-environment attestation for each run
Version or weights identifier, decoding settings, the input set, and the date and environment. It is what ties a result to one configuration rather than to the system in general.
04
A deployment-similarity argument for any performance figure relied on
MEASURE 2.3 asks for demonstration under conditions similar to the deployment setting. The artefact states the conditions a figure was produced under and the respects in which live operation differs from them.
05
A documented statement of generalisability limits
MEASURE 2.5 asks for the limits beyond the development conditions to be written down. Held by the auditee it is proof; absent, the absence is itself testable, because a supplier unable to say where a model stops working has not measured where it works.
06
A challenge-test log with adversarial cases and outcomes
Cases constructed to fail rather than drawn from ordinary traffic, with what the system did in each. Ordinary traffic exercises the middle of the distribution; controls give way at the edges, and only constructed cases go there.
§ 4 — Worked example
Worked example — forty-five referrals out of 120,000
A lender’s internal audit tests a control stated as: every application whose model score falls between 0.45 and 0.55 is referred to a human underwriter. The team selects forty-five applications across the half-year, agrees each to a referral record, finds no exceptions, and concludes the control operated effectively. Around 120,000 applications were scored in the window. The model was retrained once, in March.
What has the test actually established?
That forty-five referrals happened correctly. The referral rule is deterministic — a comparison against two constants — and it is the part of the arrangement least likely to fail; the score feeding it is the probabilistic part, and nothing here went near it. Three consequences follow. The population was available: the score sits in the decision record for all 120,000 applications, so one query would have tested the rule exhaustively for less effort than selecting the sample. That query would also have surfaced the second problem, which is that the referral rate shifts at the March retraining — the band still catches scores between 0.45 and 0.55, but not the same applications, and nobody has asked whether it still catches the borderline ones. And because the selection spans both configurations without recording which produced which score, the paper cannot separate them. The conclusion is not wrong so much as far narrower than the sentence it is written in.
§ 5 — What a weak answer looks like
A sample size taken from a table
A sample size lifted from a controls-testing table. Twenty-five items for a high-frequency automated control is a defensible number for a rule that behaves identically every time it fires, which is the assumption the table rests on. A model does not: it is one control operating a hundred thousand times with different behaviour at the margins, and the table has no view about that. The figure then gets cited in the paper as though it had settled the question of coverage, and it is the one number in the engagement nobody thinks to challenge.
§ 6 — Elsewhere
Where the declared figures come from
Where another instrument addresses the same obligation. These are correspondences, not comparisons — the Council does not rank one framework against another.
Article 15 requires a high-risk system to reach an appropriate level of accuracy, robustness and cybersecurity, with the accuracy metrics declared. Those declared figures are the population-level claim a control test is measured against.
A correspondence indicates that two instruments address the same underlying obligation. It is not a mapping endorsed by either body, not a statement that one satisfies the other, and not a judgement about which is more demanding.
§ 7 — When it applies
Which editions apply
9 January 2024
Publication date of the Global Internal Audit Standards, 2024 edition — when the current edition appeared, not a compliance deadline, and the Standards state no effective date. The NIST framework is voluntary guidance, published January 2023 and carrying no application date.
§ 8 — provenance
The provision itself
This page sets out what the instrument requires and what discharges it. The official text is the authority — these go straight to it.
- NIST AI 100-1, Artificial Intelligence Risk Management Framework 1.0 — MEASURE 2.3 and MEASURE 2.5 in the MEASURE function table.
- Global Internal Audit Standards, 2024 edition — Standard 14.1, Considerations for Implementation, on population testing and representative samples.
- Publication date printed in the Standards; no effective date is given in the document.
§ 9 — Also read
Also read
Evaluation criteria for an AI audit — IIA Standard 13.4
Certification
Assessed on the same standard of evidence
Every Council credential is examined on applied judgement against a published anchor, set out the way the obligations on this page are. The free AI Literacy Certificate is open to any adult today, and the register lists what is open for enrolment.