How to Audit AI Systems Without Missing Evidence

How to Audit AI Systems Without Missing Evidence

An auditor cannot assure an AI system that the organisation cannot identify, locate or attribute to a responsible owner. That is why how to audit ai systems starts well before a control test or technical review. It starts with a complete, governed record of what the system does, who uses it, what data it processes, and which decisions it can influence.

For compliance teams, an AI audit is not a one-off model assessment. It is a repeatable assurance process that tests whether governance requirements are operating in practice and whether the evidence would withstand scrutiny from a regulator, customer, external auditor or board committee.

An AI audit is an evidence exercise, not a model demonstration

A polished demonstration can show that an AI tool produces plausible outputs. It does not establish that the system is lawful, appropriately classified, properly monitored or controlled. An audit must test the operating environment around the model as well as the model itself.

The scope will depend on the organisation’s role. A provider developing an AI system has different obligations from a deployer using a third-party recruitment, fraud detection or customer service tool. A high-risk system under the EU AI Act demands deeper scrutiny than an internal productivity assistant with limited impact. The correct response is proportionate assurance, not the same audit template for every use case.

A useful audit asks four direct questions: is the AI system known; has it been classified correctly; are the required controls implemented; and can the organisation prove each answer with current evidence? If the answer to the final question is no, the control is not audit-ready.

How to audit AI systems: a seven-stage process

1. Set the audit boundary from the AI inventory

Begin with the AI inventory, not with a list supplied by one business unit. The audit population should include internally developed models, externally procured AI products, embedded AI features within software, pilots, proof-of-concepts and systems used by suppliers on the organisation’s behalf.

For each system, confirm the business purpose, deployment status, legal entity, jurisdictions, users, data categories, supplier, model or service version, and integrations. Record whether the system makes decisions, recommendations, rankings, predictions or generated content that people may rely on.

This step exposes a common governance failure: organisations often register the tool but not the actual use case. The same generative AI service may be low risk when drafting internal notes and materially higher risk when used to triage vulnerable customers. Audit the deployed use, not the supplier’s marketing description.

2. Confirm legal classification and applicable requirements

The next stage is to test the classification decision. Under the EU AI Act, this means determining whether the use is prohibited, high-risk, subject to transparency obligations, or outside those categories while remaining subject to other legal and governance requirements. Classification should be recorded with a rationale, supporting facts, decision owner and review date.

For high-risk use cases, assess the requirements relevant to the organisation’s role, including risk management, data governance, technical documentation, record-keeping, transparency, human oversight, accuracy, cybersecurity and post-market monitoring. Article 9 risk management and Article 17 quality management are not satisfied by a policy statement alone. The audit should identify the process, accountable owner, evidence produced and cadence of review.

Where ISO/IEC 42001 is in scope, test how the AI management system governs objectives, roles, risk treatment, operational controls, performance evaluation and continual improvement. The standard is management-system based. A collection of disconnected assessments will not demonstrate that governance is embedded.

3. Test accountability, approval and change control

Every AI system should have a named business owner and defined second-line oversight. The business owner is accountable for the use case and outcomes; legal, compliance, privacy, information security and risk functions provide challenge within their respective mandates. If ownership is shared vaguely across a programme team, it is usually owned by nobody when an issue arises.

The audit should inspect approval records for initial deployment and material changes. Materiality may include a new purpose, new data source, model replacement, altered decision threshold, expansion to another country or a change in affected individuals. Establishing a change threshold is a practical necessity. Reassessing every minor configuration adjustment creates delay; ignoring meaningful changes leaves the original assessment obsolete.

Check whether exceptions are formally approved, time-limited and tracked to closure. An informal acceptance of a missing control is not risk acceptance. It is an undocumented exposure.

4. Examine the risk assessment against the real use case

A credible assessment connects specific harms to specific controls. Generic statements such as “bias may occur” or “data protection risk is medium” are not sufficient. The audit should test whether the assessment considers affected groups, decision consequences, likelihood, severity, detectability and the limits of human intervention.

For systems affecting employment, credit, access to services, education, health or public-facing decisions, examine potential discrimination, automation bias, erroneous outputs, exclusion, manipulation and barriers to contestability. Where personal data is processed, the AI risk assessment should align with the data protection impact assessment rather than duplicate it in a separate spreadsheet.

Technical performance evidence matters, but it must be relevant. Accuracy averaged across an entire dataset may conceal poor performance for a protected group or a critical edge case. Sampling, subgroup analysis, false positive and false negative rates, and thresholds for escalation should reflect the system’s intended purpose and impact. A lower-risk content assistant may need output sampling and user guidance; a decision-support system may require far more formal validation.

5. Verify data, security and human oversight controls

The audit should trace the data lifecycle: source, lawful basis where applicable, quality checks, access permissions, retention, transfer arrangements and deletion. For third-party systems, confirm what is contractually agreed and what actually occurs. Supplier due diligence is not complete if it relies only on a questionnaire completed before procurement.

Security testing should cover access control, credential management, logging, incident response and the handling of prompt injection, data leakage, adversarial inputs or unauthorised model changes where relevant. The control depth should match the architecture. A hosted API service does not create the same testing obligations as a proprietary model, but it still requires supplier assurance and secure implementation.

Human oversight must be operational, not symbolic. Auditors should ask who can intervene, what information they receive, when they must escalate, whether they are trained to challenge outputs and whether they have sufficient authority to override the system. A human who routinely approves an AI recommendation without context is not providing meaningful oversight.

6. Inspect traceability and retained evidence

Auditability depends on records that can be retrieved without reconstructing the past from emails. For each sampled system, seek evidence of classification, risk assessment, approvals, control implementation, testing, training, monitoring, incidents, supplier reviews and change decisions.

The evidence must be versioned and attributable. An undated PDF does not prove which model version was assessed. A completed control field does not prove the control operated. Strong evidence shows the artefact, responsible person, date, scope and outcome.

Where system logging is feasible and appropriate, verify that it supports investigation of significant decisions, model behaviour and overrides. Retention periods should be documented and aligned with legal, contractual and operational requirements. Records should be protected from unauthorised alteration while remaining accessible to authorised reviewers.

7. Report findings and follow them through to closure

Classify findings according to the level of exposure and the weakness in the control environment. A missing owner, absent high-risk classification rationale or untested human oversight mechanism may be a significant finding even if no harm has yet occurred. Audits exist to identify control failure before it becomes an incident.

Each finding needs a clear statement of condition, requirement, risk, remediation action, accountable owner and due date. Avoid recommendations that merely call for “improved governance”. Specify what must be produced or changed: for example, a documented change assessment, quarterly performance review, revised user procedure or completed supplier assurance review.

Closure should require evidence, not an owner’s assurance that the action is complete. Where remediation takes time, record interim controls and the residual risk accepted by the appropriate authority.

Build an audit pack before the audit request arrives

A well-run programme keeps its evidence current rather than assembling it during a regulatory enquiry. The audit pack should be generated from the system of record and structured by AI system, framework requirement and control owner. At minimum, it should contain:

  • the current inventory and classification decisions;
  • risk, privacy, security and supplier assessments;
  • control mappings to EU AI Act and ISO/IEC 42001 requirements;
  • approvals, training records, testing results and monitoring reports; and
  • open findings, remediation status and management reporting.

This is where fragmented spreadsheets create avoidable cost. They are difficult to version, hard to challenge, and rarely provide a reliable view of ownership or evidence status. A dedicated governance platform can make the process more proportionate by connecting the inventory, legal classification, controls and audit evidence in one record rather than forcing teams to reconcile separate files.

Avoid audit theatre

Audit theatre occurs when an organisation produces policies, checklists and dashboards without testing whether they reflect operational reality. It is especially tempting with AI because terminology can obscure basic weaknesses: unknown use cases, unapproved deployments, unreviewed suppliers and untrained users.

The antidote is sampling. Select systems across risk levels, departments, suppliers and lifecycle stages. Trace a sample from inventory entry through classification, approval, implementation, monitoring and change. Then speak to the people operating the control. If the documented process says a reviewer can override an output, ask them to explain how they would do it in a live case.

The strongest AI audit is not the one with the largest evidence folder. It is the one that gives accountable leaders a clear view of where the organisation can rely on its controls, where it cannot, and what must happen next.