All research
AI Governance FrameworksBy the Helixar Research Team · July 2026 · 18 min read

Enterprise AI Governance Framework

The Helixar reference framework for governing enterprise AI systems and autonomous agents, connecting board accountability, policy, risk classification, operating controls, runtime enforcement, and audit evidence into one operating picture.

A practical, jurisdiction aware framework that turns AI principles into decision rights, controls, and evidence a board and an auditor can test.

Executive summary

  • A governance framework is only useful when it converts principles into decision rights, operating controls, and evidence that a board and an auditor can test. Statements of intent do not govern AI.
  • The Enterprise AI Governance Framework aligns with the NIST AI Risk Management Framework functions of Govern, Map, Measure, and Manage, and with the management system discipline of ISO/IEC 42001, then adds runtime enforcement and evidence design for agentic AI.
  • The framework treats agentic AI as an operational control problem, not only a model quality problem. Delegation, tool access, and autonomy define the real risk surface, and they need controls that operate while an agent is acting.
  • The framework is technology neutral and jurisdiction aware. It can carry obligations from the European Union AI Act, Australian guidance, and New Zealand privacy law without being rebuilt for each regime.
  • Every supporting report in the Helixar research roadmap connects back to this framework, and this framework links out to each of them, so an enterprise can move from the whole picture to any specific capability.

Why an enterprise needs a single AI governance framework

Most enterprises did not decide to adopt AI in one moment. It arrived through many small decisions: a team subscribed to an assistant, a vendor switched on an AI feature, a developer wired a model into a workflow, and a business unit began letting software draft, decide, or act. Each step looked minor. Together they created an estate of AI use that no single person can fully describe. A governance framework exists to bring that estate back into view and under control, so that a board can answer basic questions about what the organisation uses, what it affects, and who is accountable.

The reason a framework matters, rather than a stack of separate policies, is coherence. Without a shared structure, policy sits in one document, risk assessment in another, security review in a third, and evidence nowhere in particular. When a regulator, a customer, or the board asks how a specific AI decision was governed, the answer has to be assembled from scattered sources, and gaps appear exactly where the risk is highest. A framework gives every AI use case one path through ownership, control, and evidence, so the answer already exists.

A framework also protects innovation. Teams do not resist governance because they dislike accountability. They resist governance that is slow, unclear, or impossible to satisfy, and they respond by working around it, which pushes AI use into the shadows where risk cannot be seen. A good framework does the opposite. It gives teams a clear, proportionate path to adopt AI, so that governed adoption is easier than ungoverned adoption. The goal is not to slow AI down. It is to make the fast path also the safe path.

What the framework governs: systems, agents, and use cases

The framework governs the use of AI and the impact of the workflow, not only the ownership of a model. This distinction is the single most important design choice, because the same model can be low risk in one setting and high risk in another. A language model that drafts an internal summary is a different governance object from the same model wired into a workflow that reads customer records, calls tools, and sends communications. The unit of governance is therefore the use case, which includes the business purpose, the data, the users, the tools, the autonomy level, the human oversight, and the downstream action.

Agentic AI sharpens this point. An autonomous or semi autonomous agent does not simply answer a question. It can plan a task, choose tools, read data, write to systems, and chain actions together, sometimes without a human in the loop for each step. The risk is not only that a single output is wrong. The risk is that an agent granted excessive permissions or autonomy chains ordinary actions into a material outcome, a failure mode the OWASP Top 10 for Large Language Model Applications catalogues as excessive agency. Governing an agent means governing what it is allowed to do, not only what model sits behind it.

The framework therefore places every AI use case, whether built, bought, or switched on, into the same structure. A productivity assistant, a customer facing decision workflow, and an autonomous agent are all governed by the same layers of accountability, control, and evidence, with intensity scaled to risk. The control loop below shows how the framework treats each use case as something that moves continuously through governance rather than passing a one time review.

Framework control loop

From accountability to assurance as a repeating loop

The framework operates as a continuous loop rather than a one time assessment. A material change to model, data, workflow, autonomy, or vendor sends a use case back through the loop.

1
Govern

Board accountability, policy, risk appetite, and decision rights.

2
Map

Use case context: data, tools, delegation, and impact.

3
Measure

Testing, monitoring, and documented uncertainty.

4
Manage

Approvals, runtime enforcement, and remediation.

5
Assure

Evidence, reporting, and independent challenge.

Structure aligned to the NIST AI RMF functions and the ISO/IEC 42001 management system cycle.

The layers of the framework: from principles to evidence

The framework is built as a set of connected layers. Guiding principles set direction and encode the risk appetite of the board. Written policy turns those principles into rules for data, autonomy, oversight, and acceptable use. Lifecycle controls implement policy at design and deployment. Runtime enforcement applies policy while AI is operating. Monitoring observes behaviour in production. Evidence records what happened, and independent assurance tests the whole system. Each layer constrains the layer below it and produces evidence for the layer above, so intent at the top can be traced to action at the bottom and back again.

Making the layers explicit changes how teams work. Instead of arguing about whether the organisation has good AI governance in general, they can ask a precise question about a specific layer: whether the policy layer is clear, whether the control layer is implemented, whether the runtime layer is actually enforcing anything, and whether the evidence layer produces records an auditor could test. Weakness usually concentrates in one or two layers, and naming the layer is the first step to fixing it. The reference model report develops this architecture in detail, and the control objectives report defines what each control layer must achieve.

The layer that enterprises most often underbuild is runtime enforcement. Principles and policy are common, and design time controls are familiar from software engineering, but the ability to apply policy while an agent is acting, and to retain evidence of that enforcement, is rare. This gap matters because agents act at machine speed across many systems, and written policy alone cannot supervise them. The framework treats runtime enforcement and its evidence as first class layers, not optional extras.

Governance layers

The connected layers of the framework

Each layer constrains the one below and produces evidence for the one above. Runtime enforcement is where agentic AI risk is contained.

1
Principles and risk appetite
2
Written policy
3
Lifecycle controls
4
Runtime enforcement
5
Evidence and assurance
A technology neutral architecture that can host controls from any recognised framework.

Govern: accountability and decision rights

Govern is the function that makes the rest of the framework accountable. It answers the questions that decide whether AI governance has authority: who owns the AI portfolio, who approves use cases, who defines risk tolerance, who can accept residual risk, who trains users, who manages exceptions, who reports to the board, and who can stop a deployment. The NIST AI Risk Management Framework places these questions at the centre of its Govern function, and ISO/IEC 38507 frames them as governance responsibilities of the board and executive rather than technical choices.

Accountability works only when it connects to the structures an enterprise already trusts. AI governance should link to enterprise risk management, information security, privacy, legal, compliance, procurement, third party risk, records management, and internal audit, so that AI is governed by the organisation rather than by an isolated committee with no decision rights. The accountability model report sets out the chain from board oversight to model ownership, and the oversight models report shows how supervision should scale with risk rather than apply uniformly.

Govern also defines proportionality. Not every AI use requires the same process, and treating a low risk drafting assistant like a high impact decision system wastes effort and drives shadow AI. The framework uses risk tiers that separate low risk productivity use from high impact, safety relevant, regulated, customer affecting, or autonomous use, and the tier determines the required approvals, testing, oversight, monitoring, and evidence. Proportionality is what lets governance be strict where it matters and light where it does not.

Map and Measure: understanding and testing AI risk

Map is where the framework resists the temptation to assess AI in the abstract. Mapping makes the context of a use case visible: the business purpose, the users, the affected people, the data inputs, the retrieval sources, the tool permissions, the vendor relationships, the output audience, the decision flow, the autonomy level, and the downstream action. For agents, mapping must include delegation. Who delegates work to the agent, what identity it acts under, which tools it can call, what data it can read, what systems it can write to, and whether it can take irreversible action. These questions describe the real risk surface, and they feed the risk register.

Measure asks whether a system is fit for its intended use before deployment and whether it stays fit afterward. The relevant measures depend on the use case and its risk tier, and they extend well beyond model accuracy to include reliability, robustness, bias, privacy exposure, security behaviour, resilience to prompt injection, confabulation rate, user override rate, and operational dependency. Measurement should look at the socio technical system, not only the benchmark, because AI behaves differently once real users, real data, and business pressure enter the picture. The capability model report defines the domains this covers, and the risk register report shows how findings become tracked risk.

Documenting uncertainty is part of Measure. Leaders should know what is known, what is unknown, what cannot be measured well, and what compensating controls exist. A high impact system may carry acceptable residual uncertainty if strong human oversight, restricted autonomy, monitoring, and fallback procedures are in place. Hidden uncertainty, by contrast, becomes unmanaged risk. The framework treats the honest recording of uncertainty as a sign of governance maturity, not a weakness to conceal.

Manage and enforce: runtime control for agentic AI

Manage is where governance becomes real. Mapping and measuring identify risk, and management decides what to do about it: approve the use case, require additional controls, restrict data, reduce autonomy, require human approval, change a vendor configuration, add monitoring, block a proposed capability, accept residual risk, or retire the system. A risk assessment with no management decision is unfinished governance. Every management response should be documented, owned, and time bound where exceptions are involved.

For agents, management has to operate during use, not only at review time. A policy may state that sensitive data cannot be sent to an unapproved model, that certain tools require approval, that a high impact action must be reviewed by a human, or that unusual behaviour should trigger containment. Those requirements need a place to operate where the AI activity happens. Otherwise management relies on user memory, after the fact review, and good intentions, none of which supervise an agent acting at machine speed. Runtime enforcement does not remove the need for policy. It gives policy somewhere to act.

Management also includes containment and continuous improvement. The ability to slow, pause, or stop an agent when behaviour is unsafe is an operational control, and incident response should capture what happened, what controls operated, what failed, who was affected, and what should change. Over time, Manage should feed Govern, so the governance system learns from its own control events. This is the pattern Helixar describes as operational policy governance, and the control objectives report specifies the outcomes these controls must meet.

Assure: evidence, audit, and reporting

Assurance is what lets a board, a regulator, or a customer trust that governance is real rather than described. The framework builds assurance on evidence. During Govern, retain policy, risk appetite, roles, committee decisions, and training records. During Map, retain use case records, data maps, autonomy assessments, and impact analyses. During Measure, retain test plans, results, monitoring configuration, and override rates. During Manage, retain approvals, blocked actions, exceptions, incident reports, and residual risk acceptance. Evidence should be attributable, time stamped, access controlled, and tamper-evident where the risk justifies it.

Assurance is organised through independent lines of accountability. In the Institute of Internal Auditors Three Lines Model, the first line owns and operates controls, the second line provides risk oversight and challenge, and the third line, internal audit, provides independent assurance. External examination, such as an AICPA System and Organization Controls engagement, can add further confidence. The assurance framework report develops these lines, the evidence framework report specifies what to retain, and the audit framework report describes how internal audit tests design and operating effectiveness.

Reporting closes the loop back to the board. Good reporting is concise at the top and deep underneath, so an executive sees the position at a glance and can follow any figure down to the record behind it. The metrics and reporting framework reports set out indicators such as portfolio coverage, control effectiveness, exception ageing, and incident trends. The matrix below shows how the framework functions connect to the evidence an enterprise can review, test, and retain, which is what turns a framework into an assurance system rather than a narrative.

Framework to evidence

How framework functions map to enterprise evidence

A framework becomes governable when every high level function connects to artefacts an enterprise can review, test, and retain.

Function
Govern
Policy, risk appetite, roles, decision rights, and board reporting.
AI policy, risk tiering standard, decision rights, committee records, training records.
Map
Context, data flows, delegation, vendors, and impact.
Use case register, data map, impact assessment, tool permission map.
Measure
Performance, bias, security, reliability, and uncertainty.
Test plans, evaluation and monitoring results, override rates, residual risk notes.
Manage
Approvals, runtime control, incident response, and remediation.
Approvals, blocked actions, exceptions, incident records, risk acceptance.
Assure
Independent challenge, reporting, and board oversight.
Audit results, assurance reports, board reporting packs, remediation closure.
The unit of evidence is the use case, connected to policy, risk tier, controls, events, owners, and outcomes.

Alignment to standards and regulation

The framework is deliberately built on recognised references so that its controls are defensible and its evidence is reusable. The NIST AI Risk Management Framework supplies the operating loop of Govern, Map, Measure, and Manage, and its Generative AI Profile identifies risks such as confabulation, data leakage, and prompt injection that traditional software governance understates. ISO/IEC 42001 supplies the Plan-Do-Check-Act management system backbone, and ISO/IEC 23894 supplies AI specific risk guidance that extends ISO 31000. ISO/IEC 38507 frames the governance responsibilities that sit with the board.

On regulation, the framework is jurisdiction aware rather than jurisdiction specific. The European Union AI Act, Regulation 2024/1689, introduces obligations for high risk systems including risk management, human oversight, and record keeping, and enterprises with European exposure can carry those obligations on the same framework rather than a separate one. In Australia and New Zealand, the framework accommodates Australia’s current Guidance for AI Adoption, the Australian AI Ethics Principles, the National AI Plan approach of governing through existing law and guidance, privacy obligations, the New Zealand Algorithm Charter, and obligations under the New Zealand Privacy Act. The Australia and New Zealand focused Helixar research develops these regional obligations further.

Building on recognised references has a practical benefit beyond defensibility. An enterprise may already have a SOC 2 report or an ISO/IEC 27001 certification with controls and evidence relevant to AI governance. Reuse must be assessed against the assurance scope, criteria, systems, review period, and actual operating evidence; it is not automatic. The framework extends a relevant control environment while adding AI-specific objectives and testing where prior assurance does not apply.

Implementing the framework: a phased path

A framework is only as good as the path to adopt it, and the framework is designed to be implemented in sequence rather than all at once. Early phases build visibility and ownership, because an enterprise cannot govern what it cannot see. That means an inventory of AI use, named accountable owners, and core policy. Middle phases add risk tiering, approval workflows, and monitoring. Later phases build evidence, reporting, and independent challenge. Attempting mature assurance before basic visibility exists wastes effort, so the order matters as much as the content.

The end state is a repeatable operating capability rather than a finished project. AI use will keep changing as new models, vendors, and regulation arrive, so the goal is a governance function that can absorb change without starting over. The roadmap report sets out a phased plan across the first year, and the programme design report describes how to give the governance programme the authority, resourcing, and operating rhythm it needs to keep delivering after the initial launch energy fades.

Sequencing also manages organisational risk. A programme that tries to control everything at once tends to stall under its own weight, while a programme that starts with the highest impact and highest risk use cases delivers visible risk reduction early and earns the credibility to expand. The timeline below shows a common sequence from a fast baseline to a durable operating capability.

Implementation horizons

From baseline to operating capability

A common sequence across the first year. Early visibility enables later assurance, not the other way around.

  1. 10 to 30 days
    Baseline

    Inventory AI use, name accountable owners, and set core policy.

  2. 230 to 90 days
    Control

    Risk tier use cases, add approval workflows and monitoring.

  3. 390 to 180 days
    Assure

    Build evidence, reporting, and independent challenge.

  4. 4180 to 365 days
    Operate

    Run a repeatable capability that absorbs new use, models, and regulation.

Sequencing informed by Australia’s Guidance for AI Adoption and the NIST AI RMF.

Common failure modes the framework is designed to prevent

A framework earns its place by preventing recognisable failures, not by describing good intentions. The first failure mode is invisible AI use, where adoption outruns governance and the enterprise cannot say what it uses, what those systems touch, or who owns them. This is the most common and the most dangerous condition, because risk that cannot be seen cannot be managed. The inventory and ownership layers of the framework exist to close this gap first, which is why the phased path starts with visibility rather than sophisticated controls.

The second failure mode is governance that exists only on paper. It shows up as policy that cannot be enforced, oversight designed so that reviewers are pushed to approve rather than question, and evidence that is assembled only after an incident forces the search. Each of these looks like governance from a distance and collapses under challenge. The framework counters them by treating enforcement and evidence as operating layers, so that a policy is real when it applies while AI is acting, oversight is meaningful when a reviewer can genuinely override, and evidence exists because controls produced it rather than because someone reconstructed it later.

The third failure mode is governance that sits apart from the enterprise. When AI governance duplicates rather than connects to enterprise risk, security, privacy, legal, and audit, it competes for authority it does not have and misses the risks that cross boundaries. Agentic AI makes these boundary risks concrete, from excessive agency where chained tool calls exceed intended authority, to data leakage into unapproved models, to concentration on a single provider. The third party risk and risk register reports address these directly, and the framework connects them into the same structure so no failure mode is left in a governance blind spot.

Conclusion: the Helixar perspective

The Helixar research perspective is that an AI governance framework becomes real only when its controls can operate where AI activity happens. Principles and policy set intent, and they are necessary, but agents move faster, and across more systems, than any written document can supervise. The framework treats runtime enforcement and evidence as first class layers precisely because that is where most governance models are thin, and where agentic AI creates the exposure that boards and regulators care about most.

This is the pattern Helixar calls operational policy governance: turning written policy into rules that apply while AI is acting, observing what AI does, approving or blocking high impact actions, and retaining attributable, tamper-evident evidence that a board and an auditor can review. It is not a replacement for the NIST AI Risk Management Framework or ISO/IEC 42001. It is the operational and evidence layer that lets those frameworks be demonstrated rather than merely described.

The remaining reports in this roadmap develop each part of the framework in depth, from the capability model and assessment methodology to risk, assurance, and strategy. Read together they form a single corpus, and read individually each stands on its own. An enterprise can start anywhere its need is greatest, and the framework will connect that starting point back to the whole. For a foundational primer, the Helixar overview of enterprise AI governance is the companion to this framework.

Enterprise checklist

  • Adopt a single governance framework that connects policy, controls, and evidence rather than a stack of separate documents.
  • Govern the use case, including data, tools, autonomy, and impact, not only the model.
  • Assign accountable owners and connect AI governance to existing enterprise risk governance.
  • Match oversight, testing, and evidence to a defined risk tier.
  • Give runtime enforcement a real place to operate, so policy applies while AI is acting.
  • Retain attributable, time stamped, tamper-evident evidence for high impact AI actions.
  • Review the framework against regulatory change at least twice a year.

Frequently asked questions

How does this framework relate to the NIST AI RMF and ISO/IEC 42001?
It uses the NIST AI RMF functions of Govern, Map, Measure, and Manage as the operating loop, and the ISO/IEC 42001 management system as the accountability backbone, then adds runtime enforcement and evidence design for agentic AI. It extends these references rather than replacing them.
Is the framework specific to one regulation?
No. It is jurisdiction aware and technology neutral, so it can carry obligations from the European Union AI Act, Australia’s Guidance for AI Adoption, and New Zealand privacy law on the same structure, without being rebuilt for each regime.
What makes this framework suitable for agentic AI?
It governs the use case rather than the model alone, and it treats runtime enforcement as a first class layer. That lets an enterprise govern what an agent is allowed to do, including tool access, data, and autonomy, and retain evidence of what it actually did.
Do we need to replace our existing controls to adopt it?
No. The framework can extend an existing control environment. A SOC 2 report or ISO/IEC 27001 certification may provide relevant controls and evidence, but reuse depends on scope and operating evidence, and AI-specific objectives still need separate design and testing.
Where should an enterprise start?
Start with visibility and ownership: an inventory of AI use and named accountable owners. These reduce risk immediately and enable everything that follows. The roadmap report sets out a phased path from a baseline to a durable operating capability.

Method and source use

This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.