All research
AI AssuranceBy the Helixar Research Team · July 2026 · 19 min read

Enterprise AI Assurance Framework

How enterprises provide credible assurance that AI governance works, using the three lines model, independent challenge, and evidence, so a board can trust that controls operate rather than merely exist.

Turning governance activity into credible assurance, so trust rests on independent challenge and evidence rather than management assertion.

Executive summary

  • Assurance is the difference between claiming governance works and demonstrating it. It is what lets a board, a regulator, or a customer trust that AI is actually governed.
  • The three lines model organises assurance into who operates controls, who oversees and challenges them, and who provides independent audit, so that confidence is built rather than asserted.
  • Credible assurance depends on evidence that can be tested, not on management assertion, which is why assurance and the evidence framework are tightly linked.
  • Independent challenge is the heart of assurance. A system assured only by the team that owns it is confirmed rather than tested, which is not assurance at all.
  • For agentic AI, assurance depends on runtime evidence, because assurance of what an agent did can only rest on records of what it actually did while operating.

What assurance is and why it matters

Assurance is what turns a governance system from a set of claims into something a board and an outsider can trust. An enterprise can build policies, controls, and processes, and still be unable to demonstrate that they work, and assurance is the discipline that closes this gap. It provides credible confidence, grounded in evidence and independent challenge, that governance is not merely designed but is actually operating. Without assurance, a governance system rests on the assertion of the people who built it, which is the weakest possible foundation for trust.

Assurance matters because the people who need to trust AI governance are not the people who operate it. A board must oversee AI without running it, a regulator must judge it from outside, and a customer must rely on it without seeing it. Each of these needs a basis for trust that does not depend on believing the operators, and assurance provides that basis by having independent parties test whether governance works and report what they find. Assurance is how trust is extended to those who cannot see the governance directly.

The demand for AI assurance is growing as AI becomes more consequential and more regulated. Boards are asking whether their AI is genuinely governed, regulators are asking enterprises to demonstrate their governance, and customers and partners are asking for evidence that AI is used responsibly. Meeting these demands requires assurance that can withstand scrutiny, which is more than an internal belief that governance is sound. An enterprise that cannot provide credible assurance will increasingly find that its claims about AI governance are not accepted, however genuine they are.

Assurance tests governance

It helps to be precise about the relationship between assurance and governance, because they are often conflated. Governance is the system of accountabilities, policies, and controls that keeps AI within bounds. Assurance is the independent testing of whether that system actually works. Governance does the work; assurance checks the work. This distinction matters, because an enterprise can have strong governance and weak assurance, in which case its governance may be sound but cannot be trusted by outsiders, or weak governance and strong assurance, in which case its assurance will reveal the weakness.

Assurance is deliberately separate from governance so that it can be independent. If the same people who operate governance also assure it, the assurance is compromised, because people cannot objectively test their own work. Assurance therefore sits apart, provided by functions with enough independence to challenge governance rather than confirm it. This separation is what gives assurance its value, and an arrangement where governance assures itself, however well intentioned, provides comfort rather than assurance, because it lacks the independence that makes assurance credible.

The output of assurance feeds back into governance, closing a loop. When assurance tests governance and finds gaps, those findings drive improvements, so that governance gets better in response to independent challenge. This is the productive relationship between the two: governance operates, assurance tests, and governance improves. An enterprise where assurance findings are ignored has broken this loop, gaining the cost of assurance without the benefit, while one where findings drive improvement has assurance that makes governance stronger over time.

The three lines model

The three lines model is the established structure for organising assurance, and it applies directly to AI. In the model, as updated by the Institute of Internal Auditors, the first line owns and operates the controls, the second line provides risk oversight and challenge, and the third line, internal audit, provides independent assurance to the governing body. Each line has a distinct role and a distinct degree of independence, and together they build assurance in layers, so that confidence rests on multiple independent checks rather than a single point.

The model works because the lines are progressively more independent of the activity being assured. The first line is closest to the work and least independent, the second line has some distance and provides oversight, and the third line is fully independent and reports to the board rather than to management. This progression means that a control operated by the first line is overseen by the second and independently tested by the third, so that a weakness missed by one line may be caught by another. The layering is what makes the model robust.

The flow below shows how assurance builds across the lines up to the board. Applying the model to AI means placing AI controls clearly in the first line, giving the second line visibility and authority to challenge AI governance, and equipping internal audit to test AI controls independently. The model is not new, but AI gives it new subject matter, and an enterprise that already operates the three lines for other risks can extend it to AI rather than inventing a separate assurance structure, which is both efficient and coherent.

Three lines

How assurance builds from operation to board

Each line is progressively more independent. Evidence is what lets the second and third lines test rather than trust.

1
First line

Owns and operates AI controls.

2
Second line

Risk oversight and challenge.

3
Third line

Independent internal audit.

4
External

Independent examination where needed.

5
Board

Receives assurance and holds management to account.

Based on the IIA Three Lines Model, extended for AI specific controls and evidence.

First line: control ownership

The first line owns and operates the controls that govern AI, and it is where assurance begins, because there is nothing to assure if the controls are not operated. In the AI context, the first line includes the business and model owners who run AI use cases and the controls on them: the approvals, the data controls, the monitoring, and, for agents, the runtime enforcement. The first line responsibility is to operate these controls consistently and to produce the evidence that they operate, which the other lines will use to assure them.

The first line is the least independent of the lines, because it is doing the work being assured, and this shapes its assurance role. The first line cannot provide independent assurance of its own controls, but it provides the foundation on which the other lines build, by operating the controls well and evidencing their operation. A first line that operates controls but does not evidence them undermines the whole assurance structure, because the second and third lines have nothing to test. The first line evidence is the raw material of assurance.

Strengthening the first line is often the most effective way to improve assurance, because assurance ultimately rests on controls actually operating. An enterprise can invest heavily in second and third line assurance, but if the first line controls do not operate or do not produce evidence, the assurance will find weakness. Conversely, a strong first line that operates controls consistently and evidences them makes the work of the other lines easier and the resulting assurance stronger. Assurance is built on the first line, and a weak first line cannot be fully compensated by strong upper lines.

Second line: risk oversight and challenge

The second line provides oversight of the first line and challenge to it, sitting close enough to understand the AI controls but independent enough to question them. In the AI context, the second line typically includes the risk, compliance, and specialist AI governance functions that set the policies the first line operates and monitor whether it operates them. The second line role is not to operate the controls but to oversee them, providing a first level of independent challenge that catches problems the first line may not see in its own work.

The second line challenge is what makes it more than a monitoring function. Oversight that merely watches and reports is useful, but oversight that challenges, that questions whether controls are adequate, whether they are really operating, and whether the risks are properly understood, is what adds assurance value. A second line that confirms the first line rather than challenging it provides little assurance, because it adds no independent scrutiny. The willingness and authority to challenge is what distinguishes an effective second line from a passive one.

The second line also connects AI assurance to the enterprise risk management, because the risk functions that form much of the second line govern other risks too. This connection is valuable, because it brings AI assurance into the established risk oversight the board already trusts, rather than creating a separate AI assurance function disconnected from enterprise risk. The second line is where AI risk oversight joins the broader risk oversight, which is what makes AI assurance a normal part of enterprise assurance rather than a special case.

Third line: independent internal audit

The third line, internal audit, provides fully independent assurance to the board, and it is the strongest layer of the model because it is the most independent. Internal audit reports to the board rather than to management, which gives it the independence to challenge governance and management without fear of the people it audits. Its role is to test independently whether AI governance controls are designed properly and operate effectively, and to report its findings to the board, providing the board with an assessment it can trust because it does not come from the people who operate the governance.

Internal audit of AI governance applies the same disciplines it applies elsewhere, adapted for AI. It plans against risk, tests the design and operation of controls, and reports findings with severity and ownership. The distinction between design effectiveness and operating effectiveness is central: audit tests not only whether AI controls are well designed but whether they actually operated during the period. This is where audit depends on evidence, because operating effectiveness can only be tested from records of what controls did, which is why auditable AI governance must produce evidence.

The value of the third line to the board is that it provides an independent view, which the board cannot get from management alone. Management will report that its AI governance is sound, and it may genuinely believe so, but the board needs an independent check, and internal audit provides it. An enterprise whose board relies only on management assurance of AI governance, without independent audit, is trusting the operators to assure their own work, which is exactly the arrangement the three lines model exists to avoid. The third line independence is what makes its assurance credible to the board.

External assurance

Beyond the three internal lines, external assurance can add a further layer of confidence, provided by parties fully independent of the enterprise. External assurance is valuable where stakeholders demand it, where the enterprise wants to demonstrate its governance to customers or regulators, or where internal independence is limited. An external examination carries weight precisely because it comes from outside the enterprise entirely, and it can substantiate claims about AI governance in a way that internal assurance, however rigorous, cannot to an external audience.

Established external assurance mechanisms provide a foundation that AI assurance can build on. An AICPA System and Organization Controls examination assesses controls against defined criteria and produces a report that customers and partners can rely on, and the management system certification of ISO/IEC 42001 provides independent confirmation that an AI management system meets a recognised standard. These mechanisms extend to AI the same kind of external assurance that enterprises already provide for security and financial controls, giving external parties a basis for trust.

External assurance should build on strong internal assurance, not substitute for it. An external examination tests what the enterprise has, so an enterprise with weak internal governance and assurance will not gain strong external assurance, because the examination will find the weakness. The most effective approach is to build robust internal assurance through the three lines, producing the evidence and control operation that external assurance tests, and then to seek external examination to substantiate it to outside parties. External assurance is the capstone on internal assurance, not a replacement for it.

Evidence as the foundation of assurance

Every layer of assurance ultimately rests on evidence, because assurance is testing, and testing requires something to test. The second line challenges based on evidence, the third line audits based on evidence, and external assurance examines evidence. Without evidence that controls operate, none of the lines can do more than accept assertion, which is not assurance. This is why the assurance framework and the evidence framework are inseparable: assurance is only as strong as the evidence beneath it, and evidence exists to be used by assurance.

The strongest evidence for assurance is produced by controls as they operate, rather than assembled afterward. When a control that holds high impact actions for approval produces records of approvals and blocks as it runs, those records are evidence the assurance lines can test directly, and they are far more convincing than a retrospective account. This is why assurance is easiest in an enterprise that produces evidence as a byproduct of control operation, because the assurance lines have real records to examine rather than reconstructions to trust.

Evidence quality determines assurance quality, so the properties of good evidence matter to assurance. Evidence that is attributable, time stamped, access controlled, and tamper-evident supports strong assurance, because it can be trusted and traced. Evidence that is vague, unattributed, or alterable supports only weak assurance, because the assurance lines cannot fully rely on it. The evidence framework develops these properties, and the point for assurance is that investing in evidence quality is investing in assurance quality, because the two rise and fall together.

Assurance activities and cadence

Assurance is delivered through activities across the lines, and setting them out helps an enterprise ensure its assurance is complete rather than concentrated. The first line performs control self assessment and produces evidence. The second line performs oversight, review, and challenge. The third line performs independent audits. External parties perform examinations. Each activity has a cadence, and the cadence should follow risk, with high impact AI assured more frequently than low risk use. The matrix below sets out the assurance activities and their focus across the lines.

The cadence of assurance is a design decision that balances rigour against cost. Continuous assurance, where controls are monitored constantly and evidence is produced automatically, provides the strongest and most current assurance but requires the controls to generate evidence as they operate. Periodic assurance, where audits and examinations happen at intervals, is less current but may suffice for lower risk use. Matching the cadence to the risk, with continuous assurance for high impact and autonomous AI and periodic assurance for lower risk, keeps assurance both rigorous where it matters and feasible overall.

The activities also connect across the lines, so that assurance is a system rather than a set of isolated checks. The first line evidence feeds the second line oversight, which feeds the third line audit, which feeds the board. When these connect, assurance builds coherently from control operation to board confidence. When they are disconnected, each line works in isolation and the board receives fragmented assurance that does not add up to a coherent picture. Connecting the assurance activities across the lines is what turns them into a framework rather than a collection of activities.

Assurance activities

Assurance activity and focus across the lines

Each line contributes a distinct activity, connected so assurance builds coherently from control operation to board confidence.

Line
First line
Operate controls and self assess.
Control operation evidence.
Second line
Oversee, review, and challenge.
Oversight findings and challenge.
Third line
Independently audit design and operation.
Audit findings to the board.
External
Examine against recognised criteria.
Examination report for outside parties.
Cadence should follow risk: continuous assurance for high impact AI, periodic for lower risk.

What makes assurance credible

Not all assurance is credible, and understanding what makes it so is essential to providing it. Credible assurance rests on independence, evidence, competence, and completeness. Independence means the assurance comes from parties who can challenge rather than confirm. Evidence means the assurance tests records rather than accepting assertion. Competence means the assurers understand both governance and AI well enough to test them properly. Completeness means the assurance covers what matters rather than only what is easy to test. The stack below lists these properties of credible assurance.

The most common way assurance loses credibility is a failure of independence, where the assurance is really self assessment dressed as assurance. When the people who operate governance also assure it, the assurance confirms rather than tests, and an experienced board or regulator will recognise this and discount it. Genuine independence, uncomfortable as it is, is what makes assurance worth having, and an enterprise that seeks comfortable assurance from insiders rather than genuine assurance from independent parties is buying reassurance rather than assurance.

Credibility also depends on the assurance being willing to find problems. Assurance that never finds a weakness in a growing AI estate is more likely to be weak than to be describing perfect governance, and an assurance function under pressure to return clean results provides little value. Credible assurance is expected to find issues, treats finding them as success rather than failure, and reports them honestly. An enterprise that wants credible assurance must therefore be willing to receive uncomfortable findings, because assurance that only ever confirms is not credible however rigorous it appears.

Credible assurance

What makes assurance worth trusting

Credible assurance rests on these properties. The most common failure is a lack of genuine independence.

1
Independence: challenge, not confirmation
2
Evidence: test records, not assertions
3
Competence: understand governance and AI
4
Completeness: cover what matters
5
Willingness to find and report problems
Assurance that never finds a weakness in a growing AI estate is more likely weak than perfect.

Assurance for agentic AI

Agentic AI raises the stakes of assurance, because the question shifts from whether a model is correct to what an agent actually did, which can only be assured from records of its behaviour. Assuring an agent means being able to test what it was permitted to do, what it did, and whether it stayed within bounds, all of which requires evidence produced as the agent operated. An enterprise that cannot produce this evidence cannot assure its agents, because there is nothing for the assurance lines to test, only the design of controls that may or may not have operated.

This makes runtime evidence the foundation of agent assurance. The records of what an agent did, what it was allowed to do, and what was blocked are what let internal audit and external examination assure the agent behaviour. Without them, assurance of agents rests on testing the design of the agent controls and hoping they operated, which is design assurance without operating assurance, the weaker half of the pair. Agent assurance depends on the runtime governance that produces evidence of agent behaviour, without which it cannot be complete.

The practical implication is that an enterprise deploying agents must build the evidence that assurance requires, which means governing agents at runtime in a way that produces records. This connects assurance to operational policy governance, which enforces agent boundaries and records the result. An enterprise that governs agents this way can assure them, because the evidence exists; one that does not cannot, because no coherent record ties what the agent did to what it was permitted to do. For agentic AI, the ability to assure is a direct consequence of the ability to govern at runtime.

Common assurance failures

The most common assurance failure is the absence of genuine independence, where assurance is really self assessment by the people who operate governance. This provides comfort rather than assurance, and it fails the moment an independent party, a regulator or an external auditor, examines it. The remedy is to build the three lines properly, with real independence at the second and third lines, so that governance is challenged rather than confirmed. Independence is uncomfortable, and its discomfort is the sign that it is real.

A second failure is assurance without evidence, where the assurance lines accept management assertion because there is no evidence to test. This reduces assurance to opinion, however well informed, and it cannot substantiate governance to an outsider. The remedy is to produce evidence as controls operate, so that assurance has records to test. A third failure is assurance that never finds problems, whether because it does not look hard or because it is pressured to return clean results, which provides false comfort. The remedy is to expect and welcome findings.

A fourth failure, specific to agents, is the inability to assure what an agent did because the evidence was never captured. The remedy is to govern agents at runtime in a way that produces evidence. Each of these failures reflects a partial or compromised assurance, and the remedy in each case is to build assurance that is genuinely independent, grounded in evidence, willing to find problems, and complete enough to cover what matters, including the runtime behaviour of agents. Assurance that meets these conditions is credible; assurance that does not is reassurance.

Conclusion: the Helixar perspective

The Helixar research perspective is that assurance is only as strong as the evidence beneath it. Every layer of the three lines model, and every external examination, ultimately tests evidence, and where the evidence is weak or absent, assurance collapses into assertion. When AI controls operate through operational policy governance and produce attributable, tamper-evident records, each line of assurance can test fact rather than accept assertion, which is what makes AI assurance credible to an independent reviewer.

This matters most for agentic AI, where assurance of what an agent did can only rest on records of what it actually did. The runtime evidence that operational policy governance produces is what lets an enterprise assure its agents, testing their behaviour against their permitted authority rather than trusting that controls operated. An enterprise that governs agents at runtime can assure them; one that does not has agents it cannot fully assure, because coherent evidence tying their behaviour to their permitted authority does not exist.

Read alongside the audit, evidence, and reporting reports, this framework shows how the enterprise provides credible confidence that its AI governance works: through the three lines, independent challenge, evidence, and, where needed, external examination. Assurance is what turns governance from a set of claims into something a board and an outsider can trust, and it is the capability that makes all the other governance investment demonstrable. For the whole discipline these reports support, the Enterprise AI Governance Framework is the anchor.

Enterprise checklist

  • Separate assurance from governance so it can independently test the work.
  • Build the three lines: first line owns controls, second oversees, third audits.
  • Give the second line authority to challenge and the third line independence and board access.
  • Ground every layer of assurance in evidence produced as controls operate.
  • Consider external examination to substantiate governance to outside parties.
  • Match assurance cadence to risk, with continuous assurance for high impact AI.
  • For agents, ensure runtime evidence exists so their behaviour can be assured.

Frequently asked questions

What makes AI assurance credible?
Independence, evidence, competence, and completeness. The most important is genuine independence, so that governance is challenged rather than confirmed. Assurance based on management assertion alone is not credible to a board or regulator.
How does the three lines model apply to AI?
The first line owns and operates AI controls, the second line provides risk oversight and challenge, and the third line, internal audit, provides independent assurance to the board. External examination can add a further layer.
Do we need external assurance?
Not always, but for high impact AI or strong stakeholder demands, an independent external examination such as a SOC report or ISO/IEC 42001 certification adds confidence beyond internal audit, substantiating governance to outside parties.
Why is assurance so dependent on evidence?
Assurance is testing, and testing requires something to test. Without evidence that controls operate, the assurance lines can only accept assertion, which is not assurance. Assurance is only as strong as the evidence beneath it.
What is different about assuring agentic AI?
Assurance shifts to what an agent actually did, which can only be tested from runtime records of its behaviour. Without that evidence, agent assurance rests only on control design, the weaker half of assurance.

Method and source use

This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.