Making AI governance auditable, and auditing it well, so an independent function can test whether controls are designed and operate.
Executive summary
- Audit tests whether controls are designed well and whether they actually operated, and the gap between the two is often the most important finding.
- AI governance must be auditable by design, with controls that have clear objectives and produce evidence an auditor can test without relying on management assertion.
- Recognised control frameworks such as COBIT give auditors a defensible basis for scope and criteria, so audit conclusions are grounded rather than a matter of opinion.
- The distinction between design effectiveness and operating effectiveness is the core of audit, because a control that looks good on paper but never operated is a finding, not a pass.
- Auditing agentic AI depends on runtime evidence, because operating effectiveness for an agent can only be tested from records of what it actually did.
Source basis: ISACA COBIT; IIA Three Lines Model and position papers; AICPA SOC Suite of Services. Full citations and scope notes appear below.
Auditing AI governance
Internal audit is the independent function that tests whether AI governance actually works, and its audit of AI governance is what gives the board its most independent assurance. Audit does not operate governance or oversee it; it tests it, examining whether the controls that govern AI are designed to meet their objectives and whether they operated during the period under review. This testing, conducted independently and reported to the board, is what lets the board trust that AI governance is real rather than merely described by the management that operates it.
Auditing AI governance is both familiar and new for internal audit. It is familiar because audit has a mature discipline of testing controls, planning against risk, and reporting findings, developed over decades and codified in professional standards. It is new because AI controls have distinctive characteristics, autonomy, model behaviour, runtime enforcement, that audit must learn to test. An audit framework for AI governance adapts the established audit discipline to this new subject matter, rather than inventing a new discipline from scratch.
The value of audit is precisely its independence, which distinguishes it from the assessment and oversight that governance performs on itself. Governance functions may assess their own maturity and oversee their own controls, but only an independent audit tests them from outside, without the investment in confirming that they work. This independence is what makes audit findings credible to the board, and it is why audit sits in the third line of the assurance model, reporting to the board rather than to the management it audits.
Audit, assurance, and assessment
It helps to distinguish audit from the related activities of assurance and assessment, which are often confused. Assurance is the broad provision of confidence that governance works, delivered across the three lines. Assessment is the evaluation of governance capability, which governance may perform on itself. Audit is the specific, independent testing of controls, conducted by the third line and reported to the board. Audit is a form of assurance, the most independent form, and it uses assessment techniques, but it is distinguished by its independence and its focus on testing controls.
The relationship between audit and self assessment is particularly important, because they can look similar but differ fundamentally in independence. When a governance function assesses its own controls, it is self assessment, useful but not independent. When internal audit tests the same controls, it is audit, independent because the auditor does not operate the controls and has no investment in confirming them. The techniques may overlap, but the independence is what makes audit credible where self assessment is merely informative, and conflating the two overstates the assurance that self assessment provides.
Audit also connects to the assessment methodology, using much of the same disciplined process of scoping, evidence collection, rating, and reporting, but applying it independently. An enterprise that has a good assessment methodology has much of what audit needs, applied by an independent function. The audit framework builds on the assessment discipline, adding the independence that makes it audit, so an enterprise does not need entirely separate approaches for assessment and audit but rather the same rigorous process applied by the appropriate party, self for assessment and independent audit for audit.
Auditability by design
Audit can only test AI governance if the governance is auditable, which means auditability must be built in rather than bolted on. Auditable governance has controls with clear objectives, so the auditor knows what each control is meant to achieve and can test whether it does. And it produces evidence that an auditor can examine independently, so the auditor can test operation rather than relying on management assertion. Governance that lacks clear control objectives or produces no evidence is difficult or impossible to audit, which means its effectiveness cannot be independently verified.
Auditability by design is far cheaper than retrofitting it, because building controls to produce evidence as they operate is straightforward when done from the start and difficult when done afterward. A control designed with audit in mind records what it does as it does it, generating the evidence audit needs as a byproduct of operation. A control designed without audit in mind may operate perfectly and leave no trace, forcing audit to reconstruct what happened from indirect sources or to rely on assertion, both of which weaken the audit.
The most auditable governance is that which produces evidence automatically as controls operate, which is the runtime evidence that operational policy governance generates. When a control that enforces AI policy records each approval, block, and exception as it enforces them, audit has a complete, contemporaneous record to test, and testing operating effectiveness becomes a matter of examining the records rather than hunting for evidence. This is why auditability by design and operational governance are connected: the governance that operates controls at runtime is also the governance that is easiest to audit.
The audit process
The audit process moves through defined stages, and following them is what makes an audit rigorous and repeatable. Audit plans against risk to decide what to audit and how deeply, tests the design of controls to see whether they can meet their objectives, tests the operation of controls to see whether they actually operated, evaluates the findings and their significance, reports to the board with recommendations, and follows up to verify that agreed actions were taken. Each stage has a purpose, and skipping any of them weakens the audit.
The process is a cycle rather than a single pass, because audit is a continuing function that returns to areas over time. The follow up stage of one audit connects to the planning of the next, and areas found weak are revisited to verify improvement. This cyclical nature is what lets audit track governance over time, building a picture of whether it is improving, and it is why audit findings feed forward into future audit plans. The steps below show the audit process as a sequence that repeats.
Each stage of the process produces artefacts that document the audit and support its conclusions. Planning produces a risk based audit plan, testing produces working papers recording what was tested and found, evaluation produces the findings, reporting produces the audit report, and follow up produces the record of remediation. These artefacts are what make an audit itself auditable, allowing the audit quality to be reviewed and the audit conclusions to be traced to the evidence. An audit that does not document its process cannot demonstrate its own rigour.
From planning to follow-up
A rigorous, repeatable sequence. Follow-up of one audit connects to the planning of the next, so audit tracks governance over time.
Scope the audit against risk and define criteria.
Assess whether controls can meet their objectives.
Examine evidence that controls actually operated.
Document findings with severity, owner, and root cause.
Verify that agreed remediation actually closed the gap.
Planning against risk
Audit planning decides what to audit and how deeply, and doing this against risk is what makes audit efficient and relevant. Audit resources are finite, and auditing every AI use case to the same depth is neither possible nor sensible, so planning directs audit effort to where the risk is highest. The high impact, customer affecting, and autonomous AI use cases warrant the most audit attention, while low risk productivity use may need little. Risk based planning ensures that scarce audit effort reduces the most risk, which is the same proportionality principle that governs the rest of AI governance.
Planning also defines the criteria against which controls will be tested, and grounding these in recognised frameworks is what makes audit conclusions defensible. When audit tests AI controls against the criteria of a recognised control framework such as COBIT, the AICPA Trust Services Criteria, or the requirements of ISO/IEC 42001, its conclusions rest on an external standard rather than the auditor’s own judgement alone. This is important because audit conclusions may be challenged, and conclusions grounded in recognised criteria are far more defensible than those based on the auditor’s unaided opinion.
Planning should also draw on the enterprise risk picture, including the AI risk register, so that audit focuses on the risks the enterprise itself has identified as most significant. This connects audit to the wider governance system, so that audit is not a separate activity forming its own view of risk in isolation but an independent test of the risks the enterprise is trying to manage. An audit plan that ignores the enterprise risk register may audit the wrong things, testing controls on low risk use while high risk use goes unexamined, which is a failure of planning rather than of testing.
Testing control design
Testing control design asks whether a control is capable of meeting its objective, which is a distinct question from whether it operated. A control may be well designed, addressing its objective completely, or poorly designed, with gaps that mean it could not meet its objective even if it operated perfectly. Design testing examines the control on its merits: does it address the right risk, does it cover the relevant scenarios, and would it achieve its objective if it operated as intended. A control that fails design testing is flawed at its foundation, and its operation cannot compensate for a design that could not succeed.
For AI controls, design testing must consider the distinctive ways AI creates risk, because a control designed for traditional software may not address AI specific risks. A control designed to prevent unauthorised data access, for example, may not address the risk of data leaking into a model through a prompt, which is an AI specific pathway. Design testing for AI controls therefore asks whether the control addresses the AI specific risk it is meant to manage, not only whether it addresses the general version of that risk. A control well designed for the wrong risk is still a design failure.
Design testing is often the more revealing of the two axes for AI, because AI governance is new and controls may be designed without full understanding of the risks. An enterprise may implement a control that looks reasonable but does not actually address the AI risk, and design testing catches this before operating testing even becomes relevant, because a control that could not meet its objective by design need not be tested for operation. Strong design testing therefore catches foundational flaws that would otherwise be missed by an audit that focused only on whether controls operated.
Testing control operation
Testing control operation asks whether a control actually operated during the period, which is where audit depends most heavily on evidence. A control may be well designed and still fail to operate, whether because it was not implemented, was bypassed, or operated inconsistently, and only testing operation reveals this. Operating testing examines evidence that the control ran: records that approvals were sought and given, that prohibited actions were blocked, that monitoring occurred. Without such evidence, operating effectiveness cannot be tested, and the control operation is unknown.
The gap between design and operation is frequently the most important audit finding, because it reveals the distance between what the enterprise intends and what it does. A control that is well designed but rarely operated is a control that exists on paper and not in practice, and this gap is where governance most often fails. An audit that tested only design would report the control as sound, missing the operating gap entirely, which is why testing both axes is essential and why operating testing, though harder because it requires evidence, is indispensable.
Operating testing usually involves sampling, because controls may operate many times and audit cannot examine every instance. The auditor selects a sample of the control operations and tests whether each was performed correctly, drawing conclusions about the whole from the sample. The sampling approach must be sound and documented, so the conclusion is defensible, and it should be risk based, examining more of the high impact operations. Sampling is a strength when it is rigorous and disclosed, and a weakness when it is arbitrary or hidden, so audit treats its sampling as part of the rigour of the operating test.
Audit evidence and sampling
Audit evidence is the foundation of audit conclusions, and its quality determines the strength of the audit. Good audit evidence is sufficient, meaning there is enough of it to support the conclusion, and appropriate, meaning it is relevant and reliable. Evidence produced by controls as they operate is generally the most reliable, because it is contemporaneous and directly reflects what happened, while evidence assembled after the fact or provided by the audited party is weaker, because it is indirect or potentially self serving. Audit weighs the reliability of its evidence in forming conclusions.
The availability of good audit evidence depends on the governance being auditable by design, which is why the two connect. An enterprise whose controls produce reliable evidence as they operate gives audit strong evidence to work with, while one whose controls leave no trace forces audit to rely on weaker evidence, reconstructions and assertions, which yields weaker conclusions. The quality of an audit is therefore partly determined before the audit begins, by whether the governance was built to produce the evidence audit needs.
Sampling is how audit tests populations too large to examine entirely, and sound sampling is what makes conclusions about the whole defensible from a part. The sample should be large enough and selected appropriately, whether randomly or by risk, to support a conclusion about the population, and the approach should be documented so it can be reviewed. For AI, where controls may operate at machine speed and produce large populations of operations, sampling is essential, and risk based sampling that examines more of the high impact operations concentrates audit effort where errors would matter most.
Audit focus areas for AI
AI governance audits concentrate on a set of focus areas where the important controls sit, and setting these out helps audit ensure it covers what matters. Portfolio coverage asks whether AI use is inventoried and governed. Approvals ask whether high impact use received the required approval. Runtime control asks whether policy was enforced in operation. Incident handling asks whether incidents were detected and resolved. Each focus area pairs a control objective with the evidence audit expects to test, and together they cover the governance that most affects AI risk.
The matrix below sets out these focus areas with the evidence audit expects for each. Laying them out this way gives audit a starting point for planning and helps ensure that an audit covers the areas that matter rather than only those that are easy to test. The focus areas are not exhaustive, and an audit tailors them to the enterprise, but they represent the core of AI governance that an audit should examine, and an audit that omits any of them, particularly runtime control, leaves a significant part of AI governance unexamined.
The focus areas also reveal where auditability is often weakest, which is usually runtime control and evidence. Many enterprises can produce evidence of approvals and policies but little evidence of what actually happened at runtime, so an audit of runtime control finds that the evidence to test operating effectiveness does not exist. This is itself an important finding, because a runtime control whose operation cannot be evidenced is a control whose effectiveness is unknown, which for agentic AI is a serious gap. Audit exposing this gap is one of the more valuable things it does.
Where AI governance audits concentrate
Each focus area pairs a control objective with the evidence an auditor expects to test. Runtime control is where auditability is often weakest.
Findings, reporting, and follow-up
Audit findings record the gaps audit has identified, and their value depends on how they are expressed. A good finding states the gap clearly, cites the evidence, identifies the root cause, assesses the significance, and recommends action, so that management can understand and act on it. A finding that merely notes a problem without root cause or significance is less useful, because it does not help management fix the underlying issue or prioritise it. Audit records findings with enough detail to drive genuine improvement rather than a superficial fix.
Audit reports to the board, and the report is where audit independence delivers its value. The report gives the board an independent assessment of AI governance, which it cannot get from management, and its credibility rests on the independence and rigour of the audit behind it. The report should be clear about what was and was not audited, honest about the findings including the uncomfortable ones, and specific about their significance, so that the board can form a genuine view. An audit report that softens findings to avoid discomfort undermines the independence that gives it value.
Follow up closes the loop by verifying that agreed actions were actually taken, and it is what distinguishes audit that improves governance from audit that merely reports on it. An audit finding that is agreed but never remediated changes nothing, and follow up is how audit confirms that remediation happened and was effective. This connects audit to the continuous improvement of governance, and it is why audit tracks the status of its findings over time. Audit without follow up produces reports that are read and forgotten; audit with follow up produces improvement that can be verified.
Auditing agentic AI
Auditing agentic AI concentrates the challenge of auditability, because the operating effectiveness of an agent control can only be tested from records of what the agent actually did. An audit of an agent asks what the agent was permitted to do, what it did, and whether it stayed within bounds, all of which require runtime evidence. Where this evidence exists, the audit can test the agent operation directly, examining records of permitted, taken, and blocked actions. Where it does not, the audit can test only the design of the agent controls and must report that their operation could not be verified.
This makes runtime evidence the pivot of agent audit. An enterprise that governs agents at runtime and records their behaviour gives audit the evidence to test agent operating effectiveness, producing a complete audit. One that does not leaves audit unable to test whether the agent controls operated, which is a significant limitation, because the operating effectiveness of the controls on an agent is exactly what matters most. The auditability of agents is therefore a direct consequence of whether the enterprise governs them at runtime in a way that produces evidence.
The auditor should treat the absence of runtime evidence as a finding in its own right, not merely a limitation on the audit. When an audit cannot test whether the controls on an agent operated, because no evidence exists, this reveals that the enterprise cannot demonstrate control over its agent, which is a governance gap the audit should report. The finding is not that the agent misbehaved, which cannot be known, but that the enterprise cannot show it did not, which for a system that acts autonomously is a serious governance weakness that the audit is right to surface.
Conclusion: the Helixar perspective
The Helixar research perspective is that auditability should be built in, not bolted on. An audit can only test what the governance produces, and governance that produces no evidence of its operation cannot be audited for operating effectiveness, which is the audit that matters most. When operational policy governance records approvals, exceptions, and blocked actions as it enforces them, internal audit can test operating effectiveness directly, which shortens audits and raises confidence, because the evidence is a byproduct of governance rather than a special collection exercise.
This matters most for agents, where the operating effectiveness of controls can only be tested from runtime records. An enterprise that governs agents at runtime can be audited on what its agents actually did; one that does not can be audited only on the design of its agent controls, leaving their operation unverifiable. The ability to audit an agent is a direct consequence of the ability to govern it at runtime, which is why auditability and operational governance are two views of the same capability.
Read alongside the assurance, evidence, and control objectives reports, this framework shows how internal audit provides the board with independent assurance that AI governance works: by planning against risk, testing design and operation separately, grounding conclusions in recognised criteria, and following up to verify remediation. Audit is the independent test that turns management claims about AI governance into something the board can trust. For the whole discipline these reports support, the Enterprise AI Governance Framework is the anchor.
Enterprise checklist
- Build auditability in: controls with clear objectives that produce testable evidence.
- Plan audits against risk and ground criteria in recognised control frameworks.
- Test control design and control operation separately, and report the gap.
- Use sufficient, appropriate evidence and sound, documented sampling.
- Cover the core focus areas, including runtime control, not only the easy ones.
- Report findings with root cause and significance, and follow up to verify closure.
- For agents, test operating effectiveness from runtime evidence, and flag its absence.
Frequently asked questions
What is the difference between design and operating effectiveness?
How do we make AI governance auditable?
How does audit differ from assessment and assurance?
Why ground audit criteria in recognised frameworks?
What is special about auditing agents?
Method and source use
This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.
References
- ISACA COBIT
- IIA Three Lines Model and position papers
- AICPA SOC Suite of Services
- ISO 19011:2018, Guidelines for auditing management systems
- ISO/IEC 42001:2023, Artificial intelligence management system
- Helixar research: Enterprise AI Assurance Framework
- Helixar research: Enterprise AI Governance Framework