All research
AI Governance FrameworksBy the Helixar Research Team · July 2026 · 18 min read

Enterprise AI Governance Capability Model

A public capability model that names the governance domains an enterprise must build to manage AI, from accountability and policy to risk, assurance, evidence, and runtime control, each with observable indicators.

The EAGCM: the governance capabilities that make up a working enterprise AI governance function, and how to judge whether each one exists.

Executive summary

  • A capability model answers a different question from a maturity model. Maturity asks how well a capability operates. A capability model asks which capabilities must exist in the first place, so that nothing important is missing.
  • The Enterprise AI Governance Capability Model separates governance into observable domains, each with indicators that let an assessor judge whether the capability is absent, informal, defined, managed, or continuously assured.
  • Each domain maps to recognised references, including the NIST AI Risk Management Framework, ISO/IEC 42001, and control frameworks such as COBIT, so an assessment result is defensible rather than a matter of opinion.
  • The model is designed to be scored. It pairs with the assessment methodology and the capability assessment reports, which turn the domains into a current versus target view and a prioritised remediation plan.
  • Capability without operating evidence is only intent. The model is most valuable when each domain can be demonstrated through retained evidence, so an assessor is testing what actually happened rather than what a policy promised.

Why enterprises need a capability model

Enterprises that adopt AI quickly tend to build governance in fragments. One team writes an acceptable use policy, another sets up a review committee, a third adds monitoring to a single system, and each addition feels like progress. The problem is that no one can say whether the set of fragments adds up to a governance function or leaves large gaps. A capability model solves this by naming the complete set of governance domains an enterprise needs, so that progress can be measured against a whole rather than celebrated in parts.

The value of naming capabilities is that it makes absence visible. It is easy to see the policy that exists and hard to see the assurance capability that does not. A capability model turns silent gaps into explicit ones. When an enterprise can see that it has strong policy, moderate risk management, and almost no independent assurance, it can direct investment where the gap is real. Without the model, effort tends to flow to the capabilities that are already strong, because those are the ones people can see and improve.

A capability model also creates a shared language across functions. Risk, security, privacy, legal, engineering, and audit each tend to describe AI governance in their own terms, which makes coordination hard. When they share a model of the governance domains, they can locate their work within it, see where they depend on each other, and stop duplicating effort. This shared view is a precondition for the operating model and the programme that deliver governance in practice.

Capability and maturity: two different lenses

Capability and maturity are often confused, and keeping them distinct is essential to using either well. A capability model describes which governance capabilities must exist. A maturity model describes how well each capability operates once it does. The two are complementary. An enterprise first needs to know that a capability such as independent assurance should exist, and then needs to know whether its assurance is ad hoc, defined, managed, or continuously operating. Used together, they answer both what is missing and what is weak.

The Enterprise AI Governance Capability Model uses maturity indicators inside each domain rather than as a separate scale. For every domain it describes what the capability looks like at each level, from absent, through informal and defined, to managed and continuously assured. This lets an assessor place a domain on a scale using observable evidence rather than a general impression. The Helixar maturity model develops the levels in more depth, and the two models are designed to be used side by side.

Keeping the lenses separate also prevents a common error, which is to declare governance mature because one capability is strong. An enterprise can have a highly mature policy capability and no assurance capability at all. Maturity of a single domain says nothing about the completeness of the whole. The capability model guards against this by requiring every domain to be present before overall maturity means anything, which is why coverage of the domains comes before depth within them.

The governance domains

The model groups enterprise AI governance into a small number of domains that together cover the full discipline. The exact boundaries can be adapted, but the set should leave no material governance activity homeless. The domains are accountability and decision rights, policy management, risk management, assurance and audit, evidence and reporting, and operational and runtime control. Each is a capability an enterprise builds and matures, and each connects to the others, because governance is a system rather than a list.

Grouping matters because it determines where investment and ownership sit. A domain needs an owner, a target capability level, and a way to demonstrate that it operates. If a domain has no owner, it tends to be no one work, and it stays weak regardless of how important it is. The model therefore encourages each enterprise to assign each domain to an accountable function, so that building the capability becomes a responsibility rather than an aspiration.

The matrix below sets out the domains, what each covers, and the kind of indicator that shows the capability is present. These indicators are the basis for the assessment. They are deliberately observable, so that a rating can be supported by evidence rather than assertion, and comparable, so that results mean the same thing across business units and over time.

Capability domains

The governance domains and their indicators

Each domain is a capability an enterprise builds and matures. Indicators are observable, so a rating can be supported by evidence.

Domain
Accountability
Decision rights, ownership, and escalation for AI.
Named owners, decision rights, escalation thresholds in use.
Policy management
The lifecycle of AI policy from draft to retirement.
Owned policies with review cadence and enforcement records.
Risk management
Identifying, assessing, and treating AI risk.
Use case risk register linked to controls and owners.
Assurance and audit
Independent challenge and testing of governance.
Assurance plan, audit findings, remediation closure.
Evidence and reporting
Records and reporting that demonstrate governance.
Retained decision records and board level reporting.
Runtime control
Enforcing policy while AI is operating.
Enforcement logs, approvals, and blocked action records.
A domain with no owner tends to stay weak. The model encourages assigning each domain to an accountable function.

Accountability, policy, and risk domains

Accountability and decision rights is the domain that gives every other capability authority. A mature accountability capability names who owns the AI portfolio, who approves use cases at each risk tier, who can accept residual risk, and who can stop a deployment, and it connects those rights to existing enterprise risk governance rather than an isolated committee. The accountability model report develops this domain, including the chain from board oversight to model ownership and the escalation thresholds that move risk up the chain.

Policy management is the domain that keeps written rules current and enforceable. A mature policy capability treats policy as a lifecycle, with owners, review triggers on model, vendor, workflow, and regulatory change, and evidence that rules were applied rather than merely published. The policy management report describes this lifecycle, and it makes the point that policy which cannot be enforced creates the illusion of control while leaving the underlying risk unmanaged.

Risk management is the domain that makes AI risk visible, owned, and treated. A mature risk capability assesses risk at the use case level, captures AI specific risks such as autonomy, prompt injection, and data leakage, and links each risk to controls and evidence. The risk register report sets out the structure, and the wider risk pillar develops model risk, operational risk, third party risk, and vendor governance as specialisations of the same discipline.

Assurance, evidence, and runtime control domains

Assurance and audit is the domain that turns governance activity into credible confidence. A mature assurance capability uses independent lines of defence, so that the first line operates controls, the second line provides oversight and challenge, and the third line, internal audit, provides independent assurance, supported by evidence it can test. The assurance framework report and the audit framework report develop this domain, including the distinction between design effectiveness and operating effectiveness that a credible assessment depends on.

Evidence and reporting is the domain that lets governance be demonstrated rather than described. A mature evidence capability retains attributable, time stamped, tamper-evident records across the lifecycle, and it reports them to the board in a way that is concise at the top and deep underneath. The evidence framework report specifies what to retain, and the metrics and reporting reports describe the indicators and flows that let a board govern AI rather than receive reassurance.

Operational and runtime control is the domain that most enterprises underbuild, and it is where agentic AI creates the greatest exposure. A mature runtime capability applies policy while AI is acting, so that a high impact action can be held for approval, sensitive data cannot flow to an unapproved model, and unsafe behaviour can be contained. This is the pattern Helixar calls operational policy governance, and the control objectives report defines the outcomes these controls must meet.

Scoring and weighting the model

A capability model becomes useful when it is scored. Each domain is rated against its indicators, from absent through informal, defined, and managed, to continuously assured, and the ratings are combined into a picture of current capability. The result is not a single number but a heat map that shows where governance is strong, where it is fragile, and where it is missing. The capability assessment report describes how to produce this heat map, and the assessment methodology report describes how to gather the evidence that supports each rating.

Domains can be weighted to reflect the risk profile of the AI portfolio. An enterprise with many customer facing or autonomous systems will usually place more weight on risk, assurance, and runtime control, while an enterprise using AI mainly for internal productivity may weight policy and accountability more heavily. The weighting below is an illustrative reference model. It is a design choice an enterprise makes deliberately, not a measured statistic, and it should be reviewed as the AI portfolio changes.

Weighting should be transparent and stable. If weights change every time results are produced, the scores lose meaning and comparison over time becomes impossible. The discipline is to set weights once, based on risk, record the rationale, and change them only when the risk profile genuinely shifts. This keeps the capability model a reliable instrument rather than a flexible one that can be tuned to produce a comfortable answer.

Domain weighting

Illustrative assessment weighting across domains

A reference weighting an enterprise can adapt. Higher weight indicates a domain that carries more governance load in the portfolio being governed.

Accountability and decision rights20%
Policy management15%
Risk management20%
Assurance and audit15%
Evidence and reporting15%
Operational and runtime control15%
Illustrative reference model, not measured survey data. Set weights once from risk and change them only when the risk profile shifts.

Using the model: from assessment to investment

The purpose of a capability model is to direct investment, so the assessment should end in a plan. Once the heat map exists, an enterprise can identify the domains where the largest capability gap meets the highest risk, and treat those as priorities. A domain that is weak but low risk can wait. A domain that is weak and high risk is where the next investment belongs. This is how a capability model turns a diagnosis into a sequence of improvements rather than a report that is filed and forgotten.

The plan should be expressed as movement toward a target capability, not perfection. Not every domain needs to reach continuously assured, and pursuing the top level everywhere wastes effort. The target for each domain should reflect the risk it manages. A high target is justified for runtime control in an enterprise running autonomous agents, and a moderate target may be sufficient for a domain that manages lower impact use. Setting realistic targets keeps the programme credible and fundable.

Reassessment closes the loop. A capability model is most valuable when it is used repeatedly, so that the heat map from one cycle becomes the baseline for the next and progress is visible over time. The flow below shows the path from assessment to investment to reassessment, which is the rhythm that turns the model into a durable instrument of governance rather than a one time exercise.

Using the model

From assessment to investment to reassessment

The model is most valuable when used repeatedly, so each cycle becomes the baseline for the next.

1
Assess

Rate each domain against its indicators with evidence.

2
Prioritise

Target the largest gap that meets the highest risk.

3
Invest

Build capability toward a risk based target.

4
Reassess

Rebaseline and track progress over time.

Targets should reflect the risk a domain manages, not perfection everywhere.

The five capability levels

Within each domain, the model uses five capability levels so that a rating means the same thing wherever it is applied. At the absent level, the capability does not exist in any recognisable form, and the associated risk is simply unmanaged. At the informal level, some activity happens, but it depends on individuals, is inconsistent, and leaves little evidence. These first two levels describe most enterprises in the early period of AI adoption, when governance is a matter of good people doing their best without a system to support them.

At the defined level, the capability is documented and repeatable, with owners, procedures, and expected evidence, even if it is not yet consistently followed. At the managed level, the capability operates consistently, is measured, and produces evidence that can be tested, so that an assessor can confirm not only that it exists but that it worked during a period. The step from defined to managed is the step from having a policy to being able to prove the policy operated, and it is where many programmes stall because it requires evidence rather than intent.

At the continuously assured level, the capability operates, is measured, is independently challenged, and improves in response to its own results, so that governance learns from control events rather than repeating them. Not every domain needs to reach this level. The target for each domain should reflect the risk it manages, which keeps the model honest and the programme fundable. The timeline below shows the progression, which applies within every domain rather than to the enterprise as a whole.

Capability levels

Five levels applied within each domain

A rating means the same thing wherever it is applied. The step from defined to managed is where evidence replaces intent.

1
Level 0
Absent

The capability does not exist and the risk is unmanaged.

2
Level 1
Informal

Activity depends on individuals and leaves little evidence.

3
Level 2
Defined

Documented, repeatable, with owners and expected evidence.

4
Level 3
Managed

Operates consistently, is measured, and produces testable evidence.

5
Level 4
Continuously assured

Independently challenged and improving from its own results.

Targets should reflect the risk each domain manages, not the top level everywhere.

The runtime control domain in depth

The operational and runtime control domain deserves particular attention because it is the one enterprises most often overstate and underbuild. It is the capability to apply policy while AI is actually operating, so that a high impact action can be held for human approval, sensitive data cannot flow to an unapproved model, an agent cannot exceed its permitted tools, and unsafe behaviour can be contained before it causes harm. It is distinct from design time controls, which set a system up correctly, because it governs what the system does in production, moment to moment.

This domain matters more as autonomy increases. A recommendation system that a human reviews before acting can rely largely on oversight at the point of decision. An agent that plans, calls tools, and acts across systems cannot, because it takes many actions in the time a human takes to review one. The excessive agency failure mode, where a series of ordinary tool calls produces a material outcome, is a runtime problem, and it cannot be closed by policy documents or after the fact review alone. It needs a control that operates at the speed of the agent.

Assessing this domain honestly is difficult precisely because it is easy to claim. An enterprise may believe it controls what its agents do when its only real control is a written rule and a monthly log review. The indicator that separates belief from capability is evidence of enforcement: records of actions that were approved, held, or blocked while AI was operating. Where those records exist, the domain can be rated on fact. Where they do not, a high rating is an assertion, and the model is designed to expose exactly that gap.

Mapping the model to standards and assurance

Each domain in the model maps to recognised references, which is what makes an assessment defensible rather than a matter of internal opinion. Accountability and policy map to the Govern function of the NIST AI Risk Management Framework and to the leadership and policy clauses of ISO/IEC 42001. Risk management maps to ISO/IEC 23894 and ISO 31000. Assurance, evidence, and control objectives map to control frameworks such as COBIT and to the criteria used in AICPA System and Organization Controls examinations. Mapping the domains this way lets an enterprise cite an external standard for every capability it claims.

Mapping also lets an enterprise evaluate whether existing assurance can be reused. An organisation may have a SOC 2 report or an ISO/IEC 27001 certification with controls and evidence that overlap with governance domains such as access control and change management. Reuse is never automatic: it depends on the report or certification scope, applicable systems, review period, criteria, and operating evidence. AI-specific capabilities such as autonomy limits and model risk still require their own assessment.

The benefit of this approach is lower cost and higher credibility. A capability that is expressed in the language of a recognised standard is easier to explain to a board, an auditor, and a regulator, and it inherits the trust those standards already carry. The model therefore encourages enterprises to anchor each domain to an external reference, so that the capability model is not a private scheme but a structured view over frameworks the market already understands.

From capability model to board conversation

A capability model earns its keep when it changes what the board discusses. Boards struggle to govern AI because the subject arrives as either reassurance or alarm, and neither supports a decision. The capability model reframes the conversation around a small number of governance domains and a clear picture of which are strong, which are weak, and which are missing. That gives a board something it can act on: a view of where the organisation is exposed, and a proposal for where the next investment should go, expressed in the same language of capability and risk that governs the rest of the enterprise.

The model also gives the board a way to hold management to account over time. A single assessment is a snapshot, but a series of assessments is a trajectory, and a trajectory is what governance is really about. When the board can see that the risk management domain moved from informal to managed over a year, or that the runtime control domain has not moved at all despite growing agent use, it can ask precise questions and direct attention where it is needed. The model turns AI governance from a topic that is either fine or frightening into a programme with visible progress.

For this to work, the reporting must be honest about the weak domains, not only the strong ones. A capability model that is used to present a flattering picture loses its value immediately, because the board is then governing a fiction. The discipline is to report the heat map as it is, including the domains that are absent, and to treat those as the agenda rather than the embarrassment. Used this way, the model becomes an instrument of candour, which is what a board needs most when governing a technology that moves faster than its oversight.

Conclusion: the Helixar perspective

The Helixar research perspective is that a capability model is only as honest as the evidence beneath it. A domain can be described as mature in a slide and remain weak in reality, and the gap between the two is where governance fails quietly. The model resists this by requiring observable indicators for each domain, so that a rating is a claim about evidence rather than an impression, and by placing operational and runtime control among the core domains rather than treating it as an optional extra.

This matters most for the runtime domain, because it is the one most often overstated. An enterprise may believe it controls what its agents do, when in practice its only control is a policy document and after the fact review. Operational policy governance closes this gap by enforcing policy while AI is acting and retaining evidence of enforcement, which is what lets the runtime domain be scored on fact rather than intent.

Read alongside the framework, the assessment methodology, and the capability assessment, this model gives an enterprise a complete way to know what governance it needs, how well it operates, and where to invest next. It is the map of the discipline, and the reports it links to are the detail of each region. An enterprise that adopts the model gains a shared vocabulary, a defensible way to score itself against recognised standards, and a repeatable path from diagnosis to targeted investment. For the foundational overview, the Helixar primer on enterprise AI governance is the companion to this model.

Enterprise checklist

  • Name the complete set of governance domains so gaps become visible.
  • Assign each domain to an accountable owner and a target capability level.
  • Use capability and maturity as separate lenses: coverage first, then depth.
  • Rate each domain against observable indicators supported by evidence.
  • Weight domains by risk, set the weights once, and record the rationale.
  • Prioritise investment where the largest gap meets the highest risk.
  • Reassess on a cycle so progress is visible over time.

Frequently asked questions

How is a capability model different from a maturity model?
A capability model defines which governance capabilities must exist. A maturity model describes how well each one operates. Coverage of the domains comes first, then depth within them, and they are best used together.
How many domains should the model have?
A small number that together leave no material governance activity homeless. The Helixar model uses six: accountability, policy, risk, assurance and audit, evidence and reporting, and runtime control. The exact boundaries can be adapted.
Can the domain weighting change by industry?
Yes. The weighting is an illustrative reference model. Regulated sectors and enterprises running autonomous agents often place more weight on risk, assurance, and runtime control. Weights should be set from risk and changed only when the risk profile shifts.
How do we avoid rating a domain higher than it deserves?
Require observable evidence for each rating and separate design effectiveness from operating effectiveness, so a capability counts only if it exists and actually operated during the period.
What is the fastest way to get value from the model?
Score coverage first to find missing domains, then depth to find weak ones, and direct the next investment to the domain where the largest gap meets the highest risk.

Method and source use

This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.