All research
AI Governance FrameworksBy the Helixar Research Team · July 2026 · 19 min read

Enterprise AI Governance Reference Model

A layered reference architecture that shows how principles, policy, controls, runtime enforcement, monitoring, and evidence connect, so AI governance is coherent from board intent to operational action.

The architecture diagram for enterprise AI governance: the layers that must connect for governance to hold together, and how any control framework hosts inside them.

Executive summary

  • A reference model gives every team a shared architecture, so policy, engineering, risk, and audit describe the same system rather than talking past each other.
  • The layers move from principles and risk appetite, through policy, controls, runtime enforcement, and monitoring, to evidence and assurance, with each layer constraining the one below and producing evidence for the one above.
  • The model is technology neutral. It can host controls from the NIST AI Risk Management Framework, ISO/IEC 42001, and sector rules without being redesigned for each.
  • The layer enterprises most often underbuild is runtime enforcement, which is where agentic AI creates the most exposure and where written policy alone cannot supervise an agent acting at machine speed.
  • A reference model is not a replacement for a framework. It is the architecture that hosts a framework, so that principles at the top can be traced to enforced action and retained evidence at the bottom.

What a reference model is and why it helps

A reference model is the architecture diagram for enterprise AI governance. Where the framework describes what governance does, the reference model describes how its parts fit together as layers, from the principles that set direction at the top to the evidence that proves governance operated at the bottom. Its purpose is coherence. Without a shared architecture, teams build governance in pieces that do not connect, and the gaps appear exactly where intent at the top fails to reach action at the bottom.

The practical value of a reference model is that it gives teams a shared language and a shared picture. Policy, engineering, risk, privacy, and audit each tend to describe AI governance in their own terms, which makes coordination hard and duplication common. When they share a model of the layers, they can locate their work within it, see where they depend on one another, and stop arguing about whether governance is good in general. Instead they can ask a precise question about a specific layer, which is where improvement actually happens.

A reference model also makes weakness visible. It is easy to see that a policy exists and hard to see that the runtime layer beneath it enforces nothing. By naming every layer, the model turns silent gaps into explicit ones, so that an enterprise can see where intent is not being carried through to action. This is the same discipline the capability model applies to governance capabilities, here applied to the architecture that connects them, and the two are designed to be used together.

The layered architecture

The reference model is built from a small number of layers that together carry governance from intent to action and back. Guiding principles and risk appetite sit at the top, set by the board. Written policy turns principles into rules. Lifecycle controls implement policy at design and deployment. Runtime enforcement applies policy while AI is operating. Monitoring observes behaviour in production. Evidence records what happened, and assurance tests the whole. Each layer has a distinct job, and the coherence of governance depends on the connections between them.

The layers are ordered by abstraction, from the most general at the top to the most operational at the bottom. This ordering matters because it shows the direction of two flows. Constraint flows downward: principles constrain policy, policy constrains controls, and controls constrain what runs. Evidence flows upward: runtime enforcement produces records, records support monitoring, and monitoring feeds assurance and the board. When both flows are intact, governance is coherent, and when either breaks, intent and action drift apart.

The stack below shows the layers in order. In a healthy architecture, an enterprise can trace a single principle, such as a rule that sensitive data must not leave the organisation, down through the policy that states it, the control that implements it, the runtime enforcement that applies it, and the evidence that proves it operated. The ability to make that trace, in both directions, is the test of whether the reference model is real or only drawn.

Governance layers

The layers of the reference model

Each layer constrains the one below and produces evidence for the one above. The test is whether a single principle can be traced end to end.

1
Principles and risk appetite
2
Written policy
3
Lifecycle controls
4
Runtime enforcement
5
Monitoring and evidence
A technology neutral architecture that can host controls from any recognised framework.

Principles and policy layers

The principles layer is where the board sets direction and encodes risk appetite. Principles are broad and durable: statements about fairness, safety, accountability, privacy, and the boundaries of acceptable AI use. They change rarely, and they provide the anchor against which every lower layer is justified. The governance responsibilities that sit at this layer are the ones ISO/IEC 38507 assigns to the governing body, and they are a board matter rather than a technical one, because they express the values and risk tolerance of the organisation.

The policy layer turns principles into rules that can be applied. Where a principle says AI should respect privacy, a policy says which data may be sent to which models under what conditions. Policy is more specific than principle and more durable than a control, and it is the layer where most enterprises articulate their governance. The policy management report develops this layer, including the lifecycle that keeps policy current and the discipline that a policy is only real when it can be enforced at the layers below.

The connection between these two layers is what keeps governance principled rather than arbitrary. Every policy should trace to a principle, so that an enterprise can explain why a rule exists, and every principle should be expressed in at least one policy, so that direction does not remain an aspiration. When policies exist without principles, governance becomes a collection of rules with no rationale, and when principles exist without policies, governance becomes a statement of values with no effect. The reference model requires both, connected.

Controls and runtime layers

The controls layer implements policy at design and deployment. These are the familiar controls of software and system engineering, adapted for AI: access controls, data controls, testing, review gates, configuration, and approval workflows. They set a system up to comply with policy before it goes live, and they are necessary, but they share a limitation. A design time control governs how a system is built and configured, not what it does in operation, and for AI, and especially for agents, what a system does in operation is where the risk lives.

The runtime enforcement layer is what governs operation, and it is the layer most enterprises underbuild. It applies policy while AI is acting, so that a high impact action can be held for human approval, sensitive data cannot flow to an unapproved model, an agent cannot exceed its permitted tools, and unsafe behaviour can be contained. This layer is distinct from controls precisely because it operates continuously, moment to moment, at the speed of the AI rather than at the pace of a review. Its absence is why so much AI governance stops at intent.

The distinction between these two layers is the most important architectural point in the model. An enterprise can have strong design time controls and no runtime enforcement, which means it can build systems correctly but cannot govern what they do afterward. For assistive AI that a human reviews before acting, this may be tolerable. For agentic AI that plans, calls tools, and acts across systems, it is not, because there is no human at each step to catch a problem. The reference model treats runtime enforcement as a first class layer for exactly this reason, and the control objectives report defines what it must achieve.

Monitoring, evidence, and assurance layers

The monitoring layer observes AI behaviour in production, which is where systems reveal how they actually behave rather than how they were expected to. Monitoring watches for drift, misuse, unusual patterns, and the signals that a control is failing or a system is being pushed outside its valid range. It is the layer that turns operation into information, and it feeds both the runtime layer, which may act on what monitoring detects, and the assurance layers above, which use it to judge whether governance is working. Monitoring without action is observation, so it must connect downward as well as upward.

The evidence layer records what happened across the whole architecture, so that governance can be demonstrated rather than described. It captures the approvals, exceptions, blocked actions, incidents, and decisions that the layers below produce, and it retains them in a form that is attributable, time stamped, access controlled, and tamper-evident where risk warrants. The evidence framework report specifies what to retain, and the key architectural point is that the strongest evidence is produced by the lower layers as they operate, not assembled afterward.

The assurance layer tests the whole architecture and reports to the board. It uses independent lines of defence to challenge and verify, drawing on the evidence layer to test whether controls operated. This is where the upward flow of evidence completes, carrying the record of operational action back to the governing body that set the principles at the top. When this flow is intact, the board can govern from fact, and when it is broken, the board governs from assertion. The assurance framework report develops this layer in depth.

How the layers connect

The layers are only useful if they connect, and the connections carry the two flows that make governance coherent. The downward flow is constraint: each layer limits what the layer below may do, so that principles bound policy, policy bounds controls, and controls bound what runs. The upward flow is evidence: each layer produces records that support the layer above, so that runtime enforcement produces evidence, evidence supports monitoring and assurance, and assurance informs the board. Governance is coherent when both flows are intact end to end.

A break in either flow produces a recognisable failure. A break in the downward flow means intent does not reach action, so a principle is stated, a policy is written, and nothing enforces it, which is governance on paper. A break in the upward flow means action does not reach oversight, so controls may operate but produce no evidence, and the board cannot see whether governance worked. Most real governance failures are a break in one of these flows, and the reference model helps locate the break by making the flows explicit.

The flow below shows the connection between layers, with constraint moving down and evidence moving up. The architectural discipline is to check, for any policy that matters, that the downward flow reaches an enforcement point and the upward flow returns evidence. Where that round trip is complete, the enterprise can trust that intent becomes action and action becomes proof. Where it is not, the model shows exactly which connection to build, which is far more useful than a general instruction to improve governance.

Two flows

Constraint down, evidence up

Each layer constrains the one below and produces evidence for the one above. A break in either flow is a recognisable governance failure.

1
Principles

Constrain policy. Receive assurance.

2
Policy

Constrains controls. Reported on by evidence.

3
Controls and runtime

Constrain what runs. Produce enforcement evidence.

4
Evidence and assurance

Carry action back up to the board.

The test of the architecture is whether a policy that matters completes the round trip from intent to action to proof.

Hosting any control framework

A defining property of the reference model is that it is technology neutral and framework neutral. It does not assume one model provider, one deployment pattern, or one control catalogue. Instead it provides layers into which controls from any recognised framework can be placed. The Govern controls of the NIST AI Risk Management Framework sit at the principles and policy layers, its Map and Measure controls at the controls and monitoring layers, and its Manage controls span the runtime and evidence layers. The management system clauses of ISO/IEC 42001 map similarly.

This neutrality is valuable because it lets an enterprise adopt the architecture without abandoning the frameworks it already uses. A control catalogue such as COBIT, a security standard such as ISO/IEC 27001, and a sector rule can all be hosted in the same layered model, so that the enterprise has one architecture rather than several competing ones. The matrix below shows how common frameworks map onto the layers, which is how an enterprise checks that its chosen controls cover every layer rather than clustering at the top.

Hosting frameworks in a shared architecture also reveals coverage gaps. When an enterprise places its existing controls into the layers, it usually finds them concentrated at the principles, policy, and design control layers, with little at the runtime and evidence layers. That pattern is the architecture making a real gap visible, and it is the same gap the capability assessment finds from a different angle. The reference model and the capability model are two views of the same governance, one by architecture and one by capability, and they should agree.

Framework hosting

How recognised frameworks map onto the layers

Any control catalogue can be hosted in the layers. Placing existing controls into the model reveals where coverage clusters and where it is thin.

Layer
Principles
Direction and risk appetite.
ISO/IEC 38507 governance duties, OECD AI Principles.
Policy
Rules for data, autonomy, and oversight.
NIST AI RMF Govern, ISO/IEC 42001 policy clauses.
Controls
Design and deployment implementation.
ISO/IEC 27001 controls, COBIT, NIST Map and Measure.
Runtime
Enforcement during operation.
NIST Manage, AI control objectives, runtime policy.
Evidence
Records for assurance.
AICPA Trust Services Criteria, ISO/IEC 42001 audit.
One architecture can host many frameworks, so the enterprise avoids several competing governance stacks.

The reference model for agentic AI

Agentic AI puts the most pressure on the lower layers of the model, because an agent acts rather than only advises. For an agent, the controls layer must define the tools it may call, the data it may read, the systems it may write to, and the identity it acts under, and the runtime layer must enforce those limits while the agent operates. The risk that a series of ordinary tool calls produces a material outcome, the excessive agency failure mode, is a runtime problem, and it can only be governed at the runtime layer, not by policy documents above it.

The evidence layer also matters more for agents, because the questions that arise after an agent acts are questions of fact. What did the agent do, what did it access, what did it attempt that was blocked, and who was accountable. These can only be answered if the runtime layer produced records as the agent operated. An architecture that lacks a runtime evidence layer can govern an assistant that a human reviews, but it cannot answer for an agent, because the evidence of what the agent did was never captured.

For agentic AI, therefore, the completeness of the lower layers is not optional. An enterprise running agents needs the controls layer to bound their authority, the runtime layer to enforce those bounds, and the evidence layer to record the result, all connected so that intent reaches action and action returns proof. This is the architectural expression of what the framework calls governing the use case rather than the model, and it is why the reference model treats the runtime and evidence layers as first class rather than as extensions of design time control.

Common architecture mistakes

The most common architectural mistake is to build the top layers and neglect the bottom ones, producing an architecture that is strong on principles and policy and empty at runtime and evidence. This is comfortable to build because the top layers are documents, and it is dangerous because the layers that actually govern operation are missing. The reference model exposes this pattern immediately, because placing controls into the layers reveals the clustering at the top, and it is the single most valuable thing the model does.

A second mistake is to treat the layers as independent rather than connected. An enterprise may have activity at every layer and still fail, because the layers do not link: policy does not trace to a principle, a control does not enforce a policy, and evidence does not support assurance. Activity at each layer is necessary but not sufficient. The connections are what make governance coherent, and an architecture that has all the layers but none of the connections is a set of parts, not a system.

A third mistake is to confuse the reference model with a product or a vendor architecture. The model is technology neutral by design, and an enterprise that ties it to a single tool loses the neutrality that lets it host any framework and adapt as technology changes. The model should describe the layers and their connections, and let each enterprise implement them with whatever tools fit, so that the architecture outlasts any particular product. A reference model that depends on one vendor is not a reference model. It is a procurement decision.

Building the architecture in sequence

A reference model shows the finished architecture, but an enterprise builds it over time, and the order of construction matters. The instinct is to build top down, starting with principles and policy, because they are the most visible and the easiest to produce. That is a reasonable beginning, but it becomes a trap if it stops there, because principles and policy without the lower layers are governance that cannot act. The sequence should move deliberately down the stack, building the controls, runtime, and evidence layers that turn intent into governed action.

A more effective sequence builds enough of each layer to complete the round trip for the highest risk use cases first, rather than building each layer fully across the whole estate. For the most important policy, an enterprise builds the control that implements it, the runtime enforcement that applies it, and the evidence that proves it, so that at least one principle is governed end to end. This vertical slice through the architecture is more valuable than a complete top layer with nothing beneath it, because it proves the whole model works for the cases that matter most.

Sequencing also depends on the enterprise starting point. An organisation with strong existing security and control discipline may already have much of the controls and evidence layers and need mainly to add AI specific runtime enforcement. An organisation earlier in its governance maturity may need to build several layers. The reference model helps in both cases by showing which layers exist and which are missing, so that the enterprise builds what it lacks rather than rebuilding what it already has. The architecture guides the sequence, and the capability assessment measures the progress.

The reference model and the operating model

The reference model describes the architecture, and the operating model describes the people and decisions that run it, so the two are complementary views of the same governance. Every layer in the reference model needs owners in the operating model. The principles layer is owned by the board, the policy layer by policy owners, the controls and runtime layers by engineering and risk, and the evidence and assurance layers by risk and internal audit. A layer with no owner in the operating model tends to stay weak, however well the architecture is drawn.

Connecting the two models prevents a common disconnect, where an architecture exists on paper but no one is accountable for operating it. When each layer maps to an accountable function, the architecture becomes a set of responsibilities rather than a diagram, and the enterprise can hold specific people accountable for specific layers. This is why the reference model and the operating model report are designed to be read together, the first describing what must exist and the second describing who runs it.

The mapping also clarifies handoffs between functions, which is where governance often breaks. The controls layer is built by engineering but must implement policy owned by another function, and the evidence layer is produced by operations but consumed by assurance. Making these handoffs explicit, by mapping the architecture to the operating model, reduces the gaps that appear at the boundaries between functions. Governance fails as often at the seams between owners as within any single layer, and connecting the architecture to the operating model is how those seams are managed.

Conclusion: the Helixar perspective

The Helixar research perspective is that the runtime and evidence layers are where most reference models are thin, and where agentic AI creates the exposure that matters. Principles, policy, and design time controls are common, because they are familiar and can be expressed as documents and configuration. The ability to enforce policy while an agent is acting, and to retain evidence of that enforcement, is rare, and it is exactly the capability that operational policy governance provides. The reference model names these layers so that their absence cannot hide.

This is why the model insists on the round trip. Governance is coherent when a principle at the top can be traced down to an enforced action and back up as evidence, and it fails when that trip is broken. Most enterprises can complete the trip on paper and break it at runtime, which is precisely where operational policy governance closes the gap, by enforcing policy at the point of action and producing the evidence that carries the result back to the board.

Read alongside the framework, the capability model, and the control objectives, this reference model gives an enterprise the architecture that ties governance together, from board intent to operational action and back into assurance. It is the picture that lets teams build the right layers and connect them, rather than accumulating documents at the top and hoping governance follows. For the whole discipline these reports support, the Enterprise AI Governance Framework is the anchor.

Enterprise checklist

  • Make the governance layers explicit and shared across teams.
  • Trace every policy to a principle and every principle to a policy.
  • Map each policy that matters to the control and runtime layer that enforces it.
  • Ensure each layer produces evidence for the layer above.
  • Keep the architecture technology neutral so it can host any framework.
  • Check the round trip: intent reaches action and action returns evidence.
  • Give the runtime and evidence layers real substance, not only documents above them.

Frequently asked questions

How is a reference model different from a framework?
A framework describes what governance does. A reference model describes how its parts fit together as connected layers, from principles at the top to evidence at the bottom. The model hosts the framework rather than replacing it.
Why separate the controls layer from the runtime layer?
Design and deployment controls set a system up correctly. Runtime enforcement governs what the system does in operation, moment to moment. For agentic AI, operation is where the risk lives, so the two must be distinct.
Can the model host a framework like the NIST AI RMF?
Yes. The model is technology and framework neutral. NIST AI RMF, ISO/IEC 42001, COBIT, and sector rules all map onto the layers, so an enterprise can run one architecture rather than several competing stacks.
What does placing our controls into the model reveal?
Usually that controls cluster at the principles, policy, and design layers, with little at runtime and evidence. That pattern is the architecture making a real gap visible, the same gap a capability assessment finds from another angle.
What is the test of a healthy architecture?
Whether a single principle that matters can be traced down through policy, controls, and runtime enforcement, and back up as evidence. That complete round trip is what makes governance coherent.

Method and source use

This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.