All research
AI Governance FrameworksBy the Helixar Research Team · July 2026 · 19 min read

Enterprise AI Governance Assessment Methodology

A repeatable methodology for assessing enterprise AI governance: scoping, evidence collection, capability rating, findings, and remediation that a board and internal audit can rely on.

How to run a defensible, evidence led assessment of enterprise AI governance, end to end, so results can be trusted and retested.

Executive summary

  • A defensible assessment is evidence led rather than opinion led. It produces findings that can be retested next cycle, which is what makes governance improvement visible over time.
  • The methodology pairs with the capability model. The model defines what to assess, and the methodology defines how to assess it, so results are comparable across business units and over time.
  • Ratings should distinguish design effectiveness from operating effectiveness, in line with recognised audit and assurance practice, because a control that looks good on paper but never operated is a finding, not a pass.
  • Independence matters. An assessment run by the same team that owns the controls tends to confirm rather than challenge, so the methodology places assessment with a function that can test rather than reassure.
  • For agentic AI, the methodology extends to delegation, tool access, and runtime enforcement, because that is where the real risk surface sits and where evidence is most often missing.

Why a methodology matters

Enterprises assess AI governance all the time, but they rarely do it the same way twice, and inconsistency undermines the result. One review relies on interviews, another on a questionnaire, a third on a consultant impression, and none of them can be compared to the last. A methodology fixes this by defining a repeatable process, so that an assessment measures the same things in the same way each cycle. Repeatability is not bureaucracy. It is what turns a series of one time reviews into a trend that shows whether governance is actually improving.

A methodology also makes an assessment defensible. When a regulator, a board, or an external auditor asks how a conclusion was reached, the answer has to be more than a professional opinion. It has to be a process: this was the scope, this was the evidence, this was the rating scale, and these were the findings. A defensible methodology lets an enterprise stand behind its own assessment, and it lets a third party understand and trust the result without repeating the whole exercise.

Finally, a methodology protects against the most common failure of internal assessment, which is to grade the organisation on intent rather than operation. Without a disciplined process, assessment drifts toward confirming that policies exist and people mean well. A methodology that requires evidence and separates design from operation keeps the assessment honest, which is the only way it can be useful. An assessment that always concludes that governance is fine is not an assessment. It is reassurance.

Principles of a defensible assessment

The first principle is that assessment is evidence led. Every rating should rest on records that an independent party could examine, not on the assertion of the team being assessed. Evidence includes policy, approvals, risk assessments, monitoring output, incident records, and, for agentic AI, records of enforcement such as approved, held, and blocked actions. Where evidence does not exist, the honest rating is low, regardless of how confident the team feels. Absence of evidence is itself a finding, because governance that cannot be demonstrated cannot be relied upon.

The second principle is the separation of design effectiveness from operating effectiveness, a distinction drawn from recognised assurance practice such as AICPA System and Organization Controls engagements. Design effectiveness asks whether a control is capable of meeting its objective. Operating effectiveness asks whether it actually operated during the period under review. A control can be well designed and never operate, or operate inconsistently, and only testing both reveals the real state of governance. This distinction is the single most important discipline in the methodology.

The third principle is independence. An assessment run by the function that owns the controls will tend to confirm them, because people find it hard to challenge their own work. The methodology therefore places assessment with a second line risk function or internal audit, guided by standards such as ISO 19011 for auditing management systems and the Institute of Internal Auditors Three Lines Model. Independence does not mean adversarial. It means enough distance to test a control rather than defend it.

Phase one: scoping the assessment

Scoping decides what the assessment covers, and getting it wrong wastes the whole exercise. The scope should define the portion of the AI portfolio in range, the governance domains being assessed, the risk tiers included, and the period under review. A common mistake is to scope by system rather than by use case, which misses the context that determines risk. The methodology scopes by use case, so that the same model in a low risk and a high risk workflow is assessed as two different governance objects with different expectations.

Scoping also sets the evidence sources. Before collection begins, the assessment should identify where evidence lives: policy repositories, approval systems, risk registers, monitoring tools, incident systems, and any runtime enforcement records. Identifying sources up front prevents the assessment from quietly narrowing to whatever evidence happens to be easy to find, which biases the result toward the parts of governance that document themselves well. If a source does not exist, that too is recorded, because a missing evidence source is often the clearest sign of a missing capability.

Finally, scoping sets expectations by risk tier. A low risk productivity use case does not warrant the same assessment depth as a customer affecting decision workflow or an autonomous agent. The methodology matches assessment intensity to risk, so that scarce assessment effort concentrates where impact is highest. This proportionality keeps the assessment feasible at enterprise scale, where assessing every AI use to the same depth would be neither possible nor sensible.

Phase two: collecting evidence

Evidence collection is where the assessment either grounds itself in fact or drifts into impression. The methodology collects evidence against each capability being assessed, matching the type of evidence to the domain. For accountability, it collects decision rights, approvals, and escalation records. For policy, it collects owned policies and records that rules were applied. For risk, it collects the use case risk register and its links to controls. For runtime control, it collects enforcement records. The matrix below shows how evidence types map to what the assessment tests.

The quality of evidence matters as much as its presence. Evidence should be attributable to a person or system, time stamped, and access controlled, so that it can be trusted. A vague summary and an unstructured log both fail under scrutiny, the first because it cannot be verified and the second because it cannot be interpreted. The methodology favours evidence produced as a byproduct of control, such as an approval recorded when it was granted, over evidence assembled after the fact, which is weaker and easier to dispute.

Collection should also be sampled where volume is high. An enterprise may generate thousands of AI approvals or actions, and the assessment cannot examine every one. The methodology samples in proportion to risk, examining more of the high impact population and less of the low risk population, and it records the sampling approach so the result is reproducible. Sampling is a strength when it is disclosed and a weakness when it is hidden, so the methodology treats transparency about sampling as part of a defensible assessment.

Evidence sources

What evidence each capability requires

Evidence is matched to the domain being assessed, and favours records produced as a byproduct of control.

Capability
Accountability
Whether decision rights operated.
Approvals, escalation records, risk acceptances.
Policy
Whether rules were applied, not just published.
Owned policies, enforcement records, exceptions.
Risk
Whether risk was assessed and treated.
Use case risk register linked to controls.
Runtime control
Whether policy was enforced in operation.
Approved, held, and blocked action records.
Absence of an evidence source is itself a finding, because it often signals a missing capability.

Phase three: rating capability

Rating turns evidence into a judgement, and the methodology rates each capability on two axes rather than one. The first axis is design: is the control capable of meeting its objective as it is set up. The second axis is operation: did it actually operate during the period, consistently, and with evidence. A capability that is well designed but rarely operated is rated low on operation, and the gap between the two axes is often the most useful finding, because it shows where good intentions are not producing governed outcomes.

Ratings use the capability levels from the capability model, from absent through informal, defined, and managed, to continuously assured, applied within each domain. Using a shared scale keeps ratings comparable across the enterprise and over time, so that a rating of managed means the same thing in two business units and in two years. The methodology requires that each rating cite the evidence behind it, so a reviewer can trace the number to the record, which is what separates a defensible rating from a confident guess.

Rating should be calibrated to avoid drift. When several assessors rate different parts of the portfolio, they can apply the scale inconsistently, so the methodology includes a calibration step where assessors compare ratings against shared examples. Calibration keeps the scale stable and prevents the slow inflation that occurs when each assessor is slightly generous. A calibrated rating is one an enterprise can trust to mean the same thing every time, which is the whole point of using a scale.

Phase four: findings and severity

Findings are where the assessment becomes actionable. A finding records a gap between the target capability and the assessed capability, with enough detail that someone can act on it. The methodology records for each finding the domain, the evidence, the root cause, the risk it creates, and a severity. Severity should reflect the impact of the gap, not the ease of fixing it, so that a serious gap in a high risk domain is treated as serious even if it is inconvenient to address.

Root cause matters because it prevents the same finding from recurring. A finding that a control did not operate might have a root cause in unclear ownership, missing tooling, insufficient resourcing, or a policy that was never enforceable. Recording the root cause lets remediation fix the cause rather than the symptom, which is the difference between an assessment that improves governance and one that generates a list of repairs that reappear next cycle. The methodology treats root cause analysis as part of the finding, not an optional extra.

Findings should be proportionate and prioritised. An assessment that returns a hundred equally weighted findings overwhelms the organisation and gets ignored. The methodology ranks findings by severity and by the risk they create, so that the enterprise can act on the most important first. A small number of severe, well evidenced findings drives more improvement than a long list of minor observations, because it focuses attention where governance is genuinely exposed.

Phase five: remediation and reassessment

Remediation turns findings into change. For each finding, the methodology agrees an action, an owner, and a date, and it records residual risk where a finding cannot be closed immediately. A finding with no owner and no date is not remediation. It is a wish. The discipline of assigning ownership and time makes remediation trackable, so that the enterprise can report progress and the next assessment can verify closure rather than rediscover the same gap.

Reassessment closes the loop and is what makes the methodology repeatable. The artefacts from one cycle, the scope, the evidence, the ratings, and the findings, become the baseline for the next, so that the second assessment measures movement rather than starting from zero. This is how an enterprise builds a trajectory of governance improvement that a board can see and a regulator can trust. The flow below shows the cycle from finding through remediation to verified closure and rebaseline.

The cadence of reassessment should follow risk. High impact AI portfolios warrant more frequent assessment than low risk productivity use, and material change to a system, its data, its autonomy, or its regulatory context should trigger reassessment regardless of the calendar. The methodology treats assessment as a rhythm rather than an annual event, because AI use changes faster than a yearly cycle can track, and governance that is assessed only once a year is governance that is unmonitored for most of it.

Remediation loop

From finding to verified closure

Remediation is trackable only when each finding has an owner and a date, and closure is verified at the next cycle.

1
Finding

Gap recorded with evidence, root cause, and severity.

2
Action

Owner and date agreed, residual risk recorded.

3
Verify

Next cycle tests whether the gap actually closed.

4
Rebaseline

Updated ratings become the new baseline.

A finding with no owner and no date is not remediation.

The assessment cycle end to end

The five phases form a cycle that an enterprise runs repeatedly rather than a project it completes once. Scope defines what is assessed, collection gathers the evidence, rating produces the judgement, findings capture the gaps, and remediation drives the change, with reassessment turning the output of one cycle into the baseline for the next. Seeing the phases as a cycle keeps the enterprise focused on trajectory, which is what governance maturity actually is, rather than on the score of a single assessment.

Each phase produces artefacts that have value beyond the assessment itself. The scope becomes a record of what was and was not covered. The evidence becomes part of the assurance trail. The ratings become a heat map the board can read. The findings become a remediation plan. Because the phases produce reusable artefacts, the assessment compounds, and the second cycle is faster and richer than the first. The timeline below shows the cycle as a sequence with defined inputs and outputs.

The cycle also connects to the wider governance system. Findings feed the risk register, ratings feed board reporting, and remediation feeds the improvement backlog of the governance programme. An assessment that sits apart from these systems produces a report that is read and filed. An assessment that feeds them becomes part of how the enterprise governs, which is the outcome the methodology is designed to produce. Assessment is not the end of governance. It is one of the loops that keeps governance honest.

Assessment phases

A repeatable five phase cycle

Each phase has defined inputs, activities, and artefacts. The output of one cycle becomes the baseline for the next.

1
Phase 1
Scope

Define portfolio, domains, risk tiers, and evidence sources in range.

2
Phase 2
Collect

Gather policy, approvals, monitoring, incident, and runtime records.

3
Phase 3
Rate

Score design and operating effectiveness against the capability model.

4
Phase 4
Findings

Record gaps with evidence, root cause, and severity.

5
Phase 5
Remediate

Agree actions with owners and dates, then verify next cycle.

Phases align to recognised assurance practice and to ISO 19011 guidance on auditing management systems.

Who runs the assessment

Independence determines whether an assessment tests or reassures, so who runs it is a design decision, not an afterthought. The methodology places assessment with a function that has enough independence to challenge management and enough access to test evidence rather than accept it. In the Three Lines Model, this is typically the second line risk function for ongoing assessment and internal audit for periodic independent assurance. External assessment can add further confidence where stakeholders demand it or where internal independence is limited.

Access is as important as independence. An assessor who cannot reach the evidence is forced to rely on interviews, which returns the assessment to opinion. The methodology therefore requires that the assessing function has access to the systems where evidence lives, including approval systems, monitoring tools, and runtime enforcement records. Where access is restricted, that restriction is itself recorded as a limitation on the assessment, because an assessment that could not see the evidence is a weaker assessment and should say so.

Skill matters too. Assessing AI governance requires understanding both governance practice and the specific risks of AI, including autonomy, model behaviour, and data flows. An assessor who understands controls but not AI may miss the risks that matter most, and an assessor who understands AI but not assurance may produce findings that cannot be acted on. The methodology assumes a blend of governance and AI competence, which is increasingly the profile that second line and audit functions are building as AI use grows.

Assessing agentic AI governance

Agentic AI stretches the methodology in a specific direction, because the risk surface moves from what a model outputs to what an agent does. Assessing an agent means assessing its delegation, the identity it acts under, the tools it can call, the data it can read, the systems it can write to, and whether it can take irreversible action. These are governance questions, and they generate evidence that a traditional software or model assessment would not think to collect, particularly records of what the agent was allowed to do and what it actually did.

The evidence gap is most acute for runtime enforcement. Many enterprises can produce a policy stating what an agent may not do, but few can produce records showing that the policy was enforced while the agent was operating. The methodology treats the presence or absence of enforcement evidence as a primary indicator for agentic AI, because it is the difference between governing an agent and hoping it behaves. Where enforcement evidence exists, the runtime domain can be rated on fact. Where it does not, a high rating is unsupportable.

Assessing agents also means assessing containment. The enterprise should be able to show that it can slow, pause, or stop an agent when behaviour is unsafe, and that this capability has been tested rather than assumed. The methodology asks for evidence of containment design and, where possible, evidence of it operating, because a containment capability that has never been exercised is a claim rather than a control. This focus on operation over intent is the same discipline the methodology applies everywhere, sharpened for the case where the stakes are highest.

Common assessment pitfalls

The most common pitfall is grading intent rather than operation, which produces an assessment that always concludes governance is fine because policies exist and people mean well. The methodology guards against this by requiring evidence and separating design from operation, but the discipline has to be held, because the pressure to return a comfortable result is real. An assessment that never finds a serious gap in a growing AI estate is more likely to be weak than to be describing a strong programme.

A second pitfall is scope that quietly narrows to whatever is easy to assess. When evidence is hard to find, an assessment can drift toward the parts of governance that document themselves, leaving the harder and often riskier parts unassessed. The methodology counters this by fixing scope and evidence sources up front and by recording missing sources as findings, so that difficulty of assessment does not become a reason to ignore a domain. The riskiest domains are often the hardest to assess, which is exactly why they must not be skipped.

A third pitfall is treating the assessment as the end of governance rather than a loop within it. An assessment that produces a report which is read and filed changes nothing. The methodology connects assessment to the risk register, board reporting, and the improvement backlog, so that findings drive change and the next cycle verifies it. An assessment that does not lead to remediation and reassessment is an expense, not an investment, and the methodology is designed to make sure it is the latter.

Reporting the assessment result

An assessment only changes governance if its result reaches the people who can act on it, in a form they can use. The methodology reports the result as a heat map of capability by domain, a small number of severe findings with their remediation plans, and a trajectory that compares this cycle to the last. The heat map lets a board see the shape of governance at a glance, the findings give it decisions to make, and the trajectory tells it whether governance is improving. Reporting that is only a score, or only a list of findings, tends to be filed rather than acted upon.

Reporting should be honest about uncertainty and scope. Where the assessment sampled, it should say so. Where evidence was missing or access was limited, it should record the limitation rather than paper over it. A board that is told governance is strong, without being told what was and was not examined, is being reassured rather than informed. The methodology treats disclosure of scope and limitations as part of a credible report, because a finding that the assessment could not see something is often as important as the ratings it produced.

The report should also connect to the wider governance system so that its findings drive change. Ratings feed board reporting, findings feed the risk register and the improvement backlog, and remediation dates feed the next assessment. This is developed in the reporting framework report, which describes how assessment output becomes part of the flow of governance information from operations to the board. An assessment that produces a standalone document changes little, while one that feeds these flows becomes part of how the enterprise governs.

Conclusion: the Helixar perspective

The Helixar research perspective is that assessment quality depends on evidence quality, and evidence quality depends on how governance operates day to day. When governance produces attributable, time stamped records of approvals, exceptions, and blocked actions as controls run, an assessment becomes a test of fact rather than a review of intent. The methodology is only as strong as the evidence it can examine, and the strongest evidence is the kind that operational policy governance produces as a byproduct of enforcement.

This is why the runtime and evidence domains recur throughout the methodology. They are where assessment most often fails, not because the assessor is weak but because the evidence does not exist. An enterprise that wants defensible assessment should therefore invest in producing evidence as controls operate, so that the assessment has something to test. Evidence designed after the fact is always weaker than evidence produced in the moment, and an assessment can only be as good as the records it examines.

Read alongside the capability model and the capability assessment, this methodology gives an enterprise a complete way to judge its AI governance, defensibly and repeatably, and to improve it over time. It is the process that turns the capability model from a description into a measurement. For the whole picture of the discipline these reports support, the Enterprise AI Governance Framework is the anchor.

Enterprise checklist

  • Fix scope, domains, risk tiers, and evidence sources before collection begins.
  • Assess by use case, not by system, so context and risk are captured.
  • Collect evidence that is attributable, time stamped, and produced by controls.
  • Rate design effectiveness and operating effectiveness separately.
  • Record findings with evidence, root cause, severity, owner, and date.
  • Place assessment with a function independent enough to challenge.
  • Reassess on a risk based cadence and verify that findings actually closed.

Frequently asked questions

What makes an AI governance assessment defensible?
A repeatable process grounded in evidence: a defined scope, evidence that an independent party could examine, ratings that separate design from operating effectiveness, and findings that can be retested next cycle.
Who should run the assessment?
A function with enough independence to challenge management and enough access to test evidence, typically a second line risk function or internal audit, with external assessment where independence is limited or stakeholders demand it.
Why separate design from operating effectiveness?
A control can be well designed and never operate, or operate inconsistently. Testing both reveals the real state of governance, and the gap between them is often the most useful finding.
How often should we assess?
On a risk based cadence. High impact AI portfolios warrant more frequent assessment, and material change to a system, its data, its autonomy, or its regulatory context should trigger reassessment regardless of the calendar.
What is different about assessing agentic AI?
The risk surface moves to what an agent does: delegation, tool access, and irreversible action. The primary indicator becomes enforcement evidence, records of what the agent was allowed to do and what it actually did while operating.

Method and source use

This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.