Measuring AI governance so a board can see whether it works, using a small set of indicators that are traceable to underlying records.
Executive summary
- Metrics turn governance from a narrative into something a board can track over time, so the question ’is this working’ has an answer grounded in evidence rather than assertion.
- A useful set is small and covers coverage, risk, control effectiveness, exceptions, and incidents, because a large dashboard that no one trusts is worse than a focused set that is acted upon.
- Every metric should be traceable to underlying records, so a figure can be drilled into rather than taken on faith. A metric that cannot be traced is a story, not a measure.
- Coverage is the metric you cannot skip, because ungoverned AI use is the risk you cannot see, and coverage tells the board how much of the AI estate is actually governed.
- Metrics should mix leading indicators that warn of emerging risk with lagging indicators that record what has happened, so governance can act before an incident rather than only after.
Source basis: NIST AI RMF Playbook, Measure function; NIST AI Risk Management Framework (AI RMF 1.0); ISO/IEC 42001:2023, Artificial intelligence management system. Full citations and scope notes appear below.
Why measure governance
Governance that cannot be measured cannot be managed, and for a board, measurement is the difference between overseeing AI and being reassured about it. Principles and policies describe what should happen, but only metrics tell whether it is happening, and whether it is getting better or worse. A board that receives a narrative about AI governance is trusting the narrator, while a board that receives metrics traceable to evidence can form its own judgement. Measurement is what makes oversight independent of the confidence of the people being overseen.
Metrics also create the trajectory that governance is really about. A single snapshot of governance says where the enterprise stands today, but a sequence of measurements shows whether it is improving, stalling, or slipping, which is the more important question. Because AI use grows and changes constantly, governance is never finished, and the useful question is not whether it is perfect but whether it is keeping pace. Only metrics tracked over time can answer that, which is why measurement is a continuous capability rather than an occasional exercise.
The purpose of measurement, though, is action, not decoration. A metric exists to prompt a decision: to invest where a gap is widening, to intervene where a control is failing, or to reassure where governance is genuinely working. A metric that no one acts on is overhead, and a dashboard that is admired but never drives a decision is a cost without a benefit. The discipline of good measurement is to choose metrics that will actually change what the enterprise does, and to discard those that will not.
What good governance metrics look like
A good governance metric has a small number of properties that together make it trustworthy and useful. It is traceable, meaning the figure can be followed down to the records that produced it, so it can be verified rather than believed. It is meaningful, meaning it measures something that actually matters to risk or outcome, rather than something merely easy to count. And it is actionable, meaning a change in the metric points to a decision, rather than being a number that moves without anyone knowing what to do about it.
Traceability is the property that separates a measure from a story. A metric that reports ninety per cent coverage is only useful if someone can ask which ten per cent is not covered and get an answer from records. When a metric cannot be drilled into, it becomes an assertion dressed as a number, and a board that questions it has nowhere to go. This is why governance metrics should be built on the evidence that governance produces, so that every figure has records beneath it that can be examined.
Meaningfulness guards against the trap of measuring what is easy rather than what matters. It is easy to count how many policies exist and hard to measure whether they are enforced, so a lazy metric set reports the former and misses the latter. But the number of policies says little about whether AI is governed, while enforcement says a great deal. Good metrics resist the pull toward the countable and insist on measuring the consequential, even when the consequential is harder to capture, because a set of easy metrics can look healthy while governance is failing.
The metric families
Governance metrics fall into a small number of families that together give a rounded picture, and covering all of them prevents the blind spots that come from measuring only one. Coverage metrics show how much AI use is governed versus unknown. Risk metrics show the distribution of AI use across risk tiers and the concentration of high risk systems. Control effectiveness metrics show whether controls are operating. Exception metrics show where governance is being waived and for how long. Incident metrics show what has gone wrong and whether the enterprise is learning.
Each family answers a different question, and a set that omits a family is blind to it. An enterprise that measures control effectiveness but not coverage knows its controls work on the systems it governs but not how much AI it fails to govern at all. One that measures incidents but not exceptions sees the failures that surfaced but not the risks it has quietly accepted. Covering the families is how a metric set avoids the false comfort of looking healthy on the dimensions it happens to measure while failing on the ones it does not.
The families also connect to each other, and reading them together is more revealing than reading any alone. Rising exceptions in a domain, followed by rising incidents in the same domain, tell a story that neither metric tells by itself. Falling coverage alongside rising high risk use is a warning that governance is being outpaced. Metrics are most powerful when read as a set that tells a coherent story about the state of governance, rather than as isolated figures each judged on its own.
Coverage: the metric you cannot skip
If an enterprise could track only one governance metric, it should be coverage, because ungoverned AI use is the risk it cannot see and therefore cannot manage. Coverage measures how much of the AI estate is actually under governance: inventoried, owned, risk tiered, and subject to the appropriate controls. The gap between the AI the enterprise uses and the AI it governs is where uncontrolled risk lives, and coverage is the metric that makes that gap visible, which is why it comes first.
Coverage is harder to measure than it appears, because measuring it requires finding AI use the enterprise may not know about. Sanctioned use is easy to count, but the dangerous gap is the unsanctioned use, the AI features switched on in existing tools and the shadow use by teams under pressure. A coverage metric that counts only known use overstates coverage by ignoring exactly the use that most needs governing. Honest coverage measurement therefore includes effort to discover unknown use, and treats the discovered gap as a finding rather than a surprise.
Coverage is also the metric that most changes board behaviour, because a low coverage figure is hard to ignore. A board told that its controls are effective may be reassured, but a board told that only sixty per cent of its AI use is governed at all has a clear and uncomfortable problem to act on. Coverage reframes the governance conversation from the quality of controls on known systems to the more fundamental question of how much of the AI estate is in view, which is usually where the real exposure sits.
Control effectiveness metrics
Control effectiveness metrics answer whether the controls the enterprise relies on are actually operating, which is distinct from whether they exist. A control can be well designed and rarely operate, or operate inconsistently, and only measurement reveals the difference. Effectiveness metrics might include the proportion of high impact actions that received required approval, the rate at which prohibited actions were blocked, and the share of controls tested and found operating during a period. These measure operation, not intention, which is the distinction that matters.
These metrics depend on controls producing evidence as they operate, because effectiveness can only be measured from records of what controls did. A control that operates but leaves no record cannot be measured, so its effectiveness is unknown, which is nearly as bad as it not operating at all. This is why control effectiveness measurement pushes an enterprise toward controls that produce evidence as a byproduct, particularly at runtime, because those are the controls whose effectiveness can actually be seen rather than assumed.
Control effectiveness is also where the difference between design and operation becomes a metric. An enterprise can report that it has designed strong controls, which is a coverage of design, and separately report what proportion of them actually operated, which is coverage of operation. The gap between the two is often the most important governance metric of all, because it measures the distance between what the enterprise intends and what it does. A large gap means governance is largely on paper, however good the design looks.
Exception, remediation, and incident metrics
Exception metrics measure where governance is being waived, and they are a powerful early signal. Every exception is a deliberate acceptance of a risk the policy was meant to manage, so the volume of exceptions, their age, and their concentration by domain reveal where governance is under strain. A domain generating many exceptions is often one where the policy does not fit reality, and exceptions that outlive their expiry are risks that have quietly become permanent. Watching exceptions catches problems before they become incidents.
Remediation metrics measure whether the enterprise closes the gaps it finds. An assessment or an incident produces findings, and remediation metrics track whether those findings are actually being fixed: the proportion closed on time, the age of open findings, and the recurrence of findings that were supposedly closed. A high volume of overdue or recurring findings means the enterprise finds problems but does not fix them, which is a failure mode that looks like diligence, lots of findings, but produces no improvement.
Incident metrics measure what has gone wrong and, more importantly, whether the enterprise is learning. The count and severity of AI incidents matter, but the more revealing metrics are the time to detect, the time to contain, and whether incidents recur. Recurring incidents mean the enterprise is not learning from its own failures, which is a deeper problem than the incidents themselves. Incident metrics that focus only on the count miss the point; the question is whether the governance system gets safer as it experiences and responds to failure.
Leading and lagging indicators
Governance metrics divide into leading indicators, which warn of risk before it materialises, and lagging indicators, which record what has already happened, and a good set uses both. Incidents are lagging: they tell the enterprise about failures after they occur. Rising exceptions, growing high risk use without matching oversight, and falling control effectiveness are leading: they warn that an incident is becoming more likely. A metric set weighted toward lagging indicators tells the board about problems after they have caused harm, which is too late to prevent them.
Leading indicators are more valuable and harder to trust, which is why they are often neglected. They are more valuable because they allow prevention, and harder to trust because the link between a leading indicator and an eventual harm is probabilistic rather than certain. An enterprise that acts only on lagging indicators is always responding to yesterday failures, while one that acts on leading indicators can intervene before the failure occurs. Building a metric set that includes credible leading indicators is what shifts governance from reactive to preventive.
The best leading indicators for AI governance are often the early signs of the failure modes the enterprise most fears. For agentic AI, a rising rate of blocked actions or attempted boundary violations is a leading indicator that agents are being pushed toward their limits. A falling override rate on high impact decisions is a leading indicator that oversight is becoming a rubber stamp. Choosing leading indicators that map to real failure modes, rather than generic activity counts, is what makes them genuinely predictive rather than merely early.
The metric catalog
A metric catalog brings the families together into a defined set that the enterprise measures consistently, so that metrics mean the same thing over time and across business units. The catalog names each metric, defines how it is calculated, identifies the evidence it draws on, and states what a change in it should prompt. This discipline prevents the drift where a metric is calculated differently each period, which destroys the trajectory, and it makes each metric traceable and actionable by design rather than by chance.
The catalog should be small, because a small set that is trusted and acted upon beats a large set that is admired and ignored. Most boards can use a focused set of perhaps five to eight governance metrics, each traceable to evidence, better than a dashboard of dozens that no one has time to interrogate. The discipline of keeping the catalog small forces the enterprise to choose the metrics that matter most, which is a valuable exercise in itself, because it clarifies what the enterprise actually cares about in its AI governance.
The matrix below sets out a compact catalog across the metric families. It is a reference set that each enterprise adapts, and its value is that it covers the families without sprawling, and that each metric is defined with the evidence beneath it. A catalog built this way is what lets a board drill from a headline figure to the records that produced it, which is the property that turns a governance dashboard from a presentation into an instrument of oversight.
A compact set across the metric families
A small set that is trusted and acted upon beats a large set that is ignored. Each metric is defined with the evidence beneath it.
An illustrative KPI view
Turning the catalog into a board view means selecting the headline KPIs and setting targets, so that the board can see at a glance whether governance is where it should be. The targets express the enterprise ambition for each metric, and the gap between target and actual is what prompts action. Targets should be set from risk rather than from convenience, so that a metric like the share of high risk use cases assessed is targeted at full coverage, while a metric with less direct risk impact may carry a lower target.
The bars below show an illustrative KPI view with reference targets across the families. It is illustrative rather than measured, and each enterprise produces its own from its evidence, but the shape is instructive: coverage of high risk use and incident closure should target completeness, while broader coverage and control testing target high but not necessarily complete levels. The value of a target view is that it turns each metric into a clear statement of whether the enterprise is meeting its own governance ambition.
A KPI view is only as good as the evidence beneath it, which is the theme that runs through all of governance measurement. A target of ninety per cent coverage is meaningful only if the coverage figure is traceable to records of what is and is not governed. A KPI view built on figures that cannot be drilled into is a presentation, not an instrument, and a board that senses this will not trust it. The discipline of traceability is what lets a KPI view survive the scrutiny of a board that questions a number.
Illustrative governance KPI targets
Reference targets across the families. Coverage of high risk use and incident closure target completeness; broader measures target high levels.
Metrics for agentic AI
Agentic AI needs metrics that traditional software governance does not, because the risk surface is different. For agents, useful metrics include the rate of actions blocked at the runtime boundary, the frequency of attempted actions outside permitted scope, the proportion of high impact agent actions that required and received approval, and the time to contain an agent behaving unsafely. These measure the specific ways agents create risk, and they are only available if the enterprise enforces and records agent behaviour at runtime.
These agent metrics are among the most valuable leading indicators an enterprise can have, because they show pressure on the boundaries before a boundary is breached. A rising rate of blocked or attempted out of scope actions suggests that agents are being pushed toward their limits, whether by design changes, prompt manipulation, or expanding use, and it warns of a potential failure before one occurs. An enterprise that watches these metrics can tighten limits or investigate before an agent causes harm, which is prevention rather than response.
The availability of agent metrics is itself a governance signal. An enterprise that cannot produce them is one that does not enforce or record agent behaviour at runtime, which means its oversight of agents rests on design and hope rather than observation. The presence of rich agent metrics, by contrast, indicates an enterprise that governs its agents where they act. In this sense the metrics are not only a measure of agent behaviour but a measure of whether the enterprise has the runtime governance that agentic AI requires.
From metrics to decisions, and avoiding vanity metrics
The point of governance metrics is to drive decisions, and the test of a metric set is whether it changes what the enterprise does. A metric that reveals a widening coverage gap should prompt investment in discovery and onboarding. A falling override rate should prompt a review of whether oversight has become a rubber stamp. Rising exceptions in a domain should prompt a look at whether the policy fits reality. Each metric should have a defined response, so that measurement leads to action rather than to a dashboard that is watched but not used.
The great danger is the vanity metric, a figure chosen because it looks good rather than because it reveals the truth. Counting the number of policies, the hours of training delivered, or the number of committee meetings held all produce impressive numbers that say little about whether AI is governed. Vanity metrics are seductive because they are easy to improve and comfortable to report, but they measure activity rather than outcome, and a metric set full of them can look healthy while governance fails. The stack below lists common pitfalls to avoid.
Guarding against vanity metrics requires asking, for each metric, what decision it would change and what it would reveal that the enterprise would rather not know. A metric that can only ever report good news is probably a vanity metric. A metric that occasionally forces an uncomfortable conversation is probably a real one. The willingness to measure things that might look bad, like coverage gaps and near zero override rates, is what separates a metric set that governs from one that reassures, and it is the discipline that keeps measurement honest.
Common measurement failures to avoid
Each pitfall produces numbers that look healthy while governance may be failing. Prefer metrics that can force an uncomfortable conversation.
Conclusion: the Helixar perspective
The Helixar research perspective is that a governance metric is only as good as the evidence beneath it. A figure that cannot be drilled into is a story, and a board that governs from stories is governing blind. When coverage, control effectiveness, exceptions, and incidents are all traceable to records that governance produced as it operated, each metric becomes a summary of fact that the board can interrogate, and measurement becomes an instrument of oversight rather than a presentation.
This is why measurement depends on the evidence that operational policy governance produces. When approvals, blocked actions, exceptions, and control events are recorded as they happen, the metrics that summarise them are trustworthy by construction, because the records beneath them are real. An enterprise that enforces and records governance at runtime finds that its metrics largely build themselves, because the evidence they need is a byproduct of the controls operating, rather than a special collection effort.
Read alongside the reporting and assessment reports, this report shows how the enterprise knows whether its AI governance works: through a small, traceable set of metrics across the families, mixing leading and lagging indicators, and always answerable down to the evidence. Metrics are how governance becomes visible to the people accountable for it, and they are the difference between overseeing AI and being reassured about it. Chosen well, kept small, and grounded in evidence, they let a board watch its AI governance improve or slip with the same clarity it brings to its finances. For the whole discipline these reports support, the Enterprise AI Governance Framework is the anchor.
Enterprise checklist
- Choose a small set of metrics the board will actually use and act upon.
- Cover the families: coverage, risk, control effectiveness, exceptions, incidents.
- Make every metric traceable to the underlying evidence it draws on.
- Measure coverage first, because ungoverned AI is the risk you cannot see.
- Mix leading indicators that warn with lagging indicators that record.
- Add agent specific metrics such as blocked actions and time to contain.
- Define a response for each metric, and reject vanity metrics that only look good.
Frequently asked questions
What is the single most useful governance metric?
How many KPIs should a board see?
What is the difference between leading and lagging indicators?
What metrics matter for agentic AI?
How do we avoid vanity metrics?
Method and source use
This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.
References
- NIST AI RMF Playbook, Measure function
- NIST AI Risk Management Framework (AI RMF 1.0)
- ISO/IEC 42001:2023, Artificial intelligence management system
- COSO Enterprise Risk Management Framework
- Helixar research: AI Governance Reporting Framework
- Helixar research: Enterprise AI Governance Capability Assessment
- Helixar research: Enterprise AI Governance Framework