Turning the capability model and assessment methodology into a scored current versus target view and a plan the enterprise can act on.
Executive summary
- A capability assessment converts the capability model into a current state score, a target state, and a prioritised gap, so an enterprise knows not only what governance it has but where to invest next.
- Scoring should be evidence based and comparable across business units and over time, so that a rating means the same thing everywhere and a trajectory of improvement becomes visible.
- The output is a heat map that a board can read at a glance and a remediation roadmap that links the largest gaps to the highest risk, not a single vanity score.
- Target state should be set by risk, not by ambition. Not every domain needs the top capability level, and pursuing it everywhere wastes effort and credibility.
- The assessment is credible only when each score is supported by retained evidence, so a self reported rating without evidence is treated as a finding rather than a result.
Source basis: NIST AI Risk Management Framework (AI RMF 1.0); NIST AI RMF Playbook; ISO/IEC 42001:2023, Artificial intelligence management system. Full citations and scope notes appear below.
What a capability assessment produces
A capability assessment produces three things: a score for each governance domain, a target for each domain, and a prioritised plan to close the gap between them. The score describes the current state, the target describes where the domain needs to be given its risk, and the gap describes the work. Together they turn a general sense that AI governance needs attention into a specific, fundable programme. Without an assessment, investment tends to flow to whatever is visible or fashionable rather than to where the exposure is greatest.
The most useful form of the output is a heat map, because it communicates the shape of governance at a glance. A board can see in one view which domains are strong, which are fragile, and which are missing, and it can direct attention accordingly. A single overall score, by contrast, hides more than it reveals, because it averages a strong policy capability and an absent assurance capability into a middling number that suggests neither the strength nor the danger. The assessment favours the heat map over the average for exactly this reason.
The assessment also produces a trajectory when it is repeated. A single assessment is a snapshot, but a sequence of assessments shows whether governance is improving, stalling, or slipping, which is what a board most needs to know. The value of the assessment therefore compounds with repetition, and an enterprise that assesses once learns far less than one that assesses on a cycle and watches the heat map change. Assessment is a rhythm, not an event.
Inputs: the model and the methodology
A capability assessment does not stand alone. It applies the assessment methodology to the capability model, so the two upstream reports are its inputs. The capability model provides the domains and the indicators that define what good looks like. The methodology provides the disciplined process for gathering evidence, rating design and operating effectiveness, and recording findings. The assessment is the point where these come together to produce a scored result, which is why it should never be attempted without both a defined model and a defined method.
Using a shared model and method is what makes assessments comparable. If two business units score themselves against different domains using different processes, their results cannot be compared, and the enterprise cannot see its overall position. When both use the same capability model and the same methodology, a rating of managed in one unit means the same as a rating of managed in another, and the enterprise can aggregate results into a portfolio view. Comparability is not a nicety. It is what lets an enterprise govern AI at scale rather than one team at a time.
The inputs also include the risk profile of the AI portfolio, because that is what sets the target. A domain target is not a fixed ideal. It is a function of the risk the domain manages, so an enterprise running autonomous agents needs a higher target for runtime control than one using AI only for internal drafting. The assessment therefore begins by understanding the portfolio and its risk, so that the targets it sets are grounded in exposure rather than in a generic notion of best practice.
Scoring current state
Scoring current state means rating each domain against its indicators using evidence. The assessment applies the five capability levels, from absent through informal, defined, and managed, to continuously assured, and it requires that each rating cite the evidence behind it. A domain is not scored on how confident its owner feels but on what the records show, which is why the methodology insists on collecting attributable, time stamped evidence before rating begins. The score is a claim about evidence, and it should be traceable to the record that supports it.
Current state scoring should separate the two axes of design and operation, because they often differ. A domain can be well designed, with clear procedures and owners, and still operate poorly, with controls that rarely run or run inconsistently. Recording both axes shows where good design is not producing governed outcomes, which is frequently the most actionable insight the assessment provides. A domain rated well on design and poorly on operation is a domain where the fix is execution, not redesign, and knowing that saves wasted effort.
Scoring should also be calibrated across assessors to prevent drift. When several people rate different parts of the portfolio, they can apply the scale slightly differently, and small inconsistencies accumulate into an unreliable picture. A calibration step, where assessors compare their ratings against shared examples, keeps the scale stable so that a rating means the same thing regardless of who produced it. Calibration is what allows an enterprise to trust its own heat map, and without it the scores slowly lose meaning.
Setting target state
Target state is the level each domain needs to reach, and setting it well is what keeps the programme both credible and affordable. The temptation is to set every domain to the top level, but that is neither necessary nor wise, because the top level is expensive to build and maintain and is only justified where the risk warrants it. The assessment sets targets by risk, so a domain that manages high impact or autonomous AI is given a high target, while a domain managing low impact use is given a moderate one.
Targets should be explicit and stable. If the target moves each time the assessment runs, the gap becomes meaningless and comparison over time breaks down. The assessment records the target and the rationale for it, and it changes the target only when the risk profile genuinely shifts, such as when the enterprise begins running autonomous agents where it previously used only assistive AI. Stable targets make the gap a reliable measure of progress rather than a moving line that can be redrawn to flatter the result.
Setting targets is also a governance decision, not a technical one, so it belongs with accountable owners rather than assessors alone. The people accountable for AI risk should agree what capability each domain needs, because they are the ones who will fund the gap and answer for the residual risk. The assessment facilitates this decision by showing what each target level requires and what risk a lower target leaves unmanaged, so that the choice is informed. Targets set without accountable ownership tend to be ignored when funding is decided.
The capability gap and the heat map
The gap between current and target state is the core output of the assessment, and the heat map is how it is communicated. For each domain, the heat map shows the current level, the target level, and the distance between them, so that the pattern of exposure is visible at a glance. A domain at target is green, a domain a level below is amber, and a domain well below target or absent is red. The colour is a summary, and behind each cell sits the evidence and the findings that justify it.
The heat map is deliberately honest about weakness, which is what gives it value. A heat map that shows every domain as strong is either describing an unusually mature enterprise or, far more often, hiding the truth to avoid discomfort. The assessment reports the gaps as they are, including the domains that are absent, and treats those as the agenda rather than the embarrassment. A board governing from an honest heat map can act. A board governing from a flattering one is governing a fiction.
The heat map below shows the structure, with each domain rated against its target and the gap made explicit. In practice the enterprise populates it with its own evidence based ratings, and the pattern usually tells a story: policy and accountability tend to be stronger because they document themselves, while assurance, evidence, and runtime control tend to lag because they require more than a document to exist. The heat map turns that story into something a board can act on.
Current against target across domains
The gap between current and target is the core output. Each cell is backed by evidence and findings, not impression.
Weighting and aggregation
While the heat map is the primary output, an enterprise sometimes needs an aggregate view, and weighting is how the domains combine into one. Weighting reflects how much governance load each domain carries for the specific AI portfolio, so risk, assurance, and runtime control usually carry more weight where autonomy is high. The weighting is an illustrative reference model, a deliberate design choice rather than a measured statistic, and it should be set once from risk and left stable so the aggregate means something over time.
Aggregation should be used with care, because a single number hides the pattern the heat map reveals. The assessment produces an aggregate only as a secondary summary, always alongside the heat map, never instead of it. An aggregate score is useful for tracking overall trajectory across cycles, but it should never be the basis for a decision, because a decision needs to know which domain is weak, not just that the average is middling. The donut below shows an illustrative weighting across the assessment dimensions.
When an aggregate is reported, the weighting behind it must be disclosed. A board reading a single governance score needs to know that it reflects a particular weighting choice, so that it can judge whether the emphasis is right for the enterprise. Undisclosed weighting turns an aggregate into a black box, and a black box is not something a board can govern from. Transparency about weighting is part of what keeps the assessment defensible rather than a number that has to be taken on trust.
Illustrative weighting of assessment dimensions
A reference weighting across the dimensions of the assessment. Enterprises calibrate this to their own AI risk exposure.
- Govern and accountability25%
- Map and risk20%
- Measure and monitor20%
- Manage and control20%
- Assure and evidence15%
From heat map to remediation roadmap
A heat map that leads nowhere is a diagnosis without a treatment. The assessment turns the heat map into a remediation roadmap by taking each material gap and defining the work to close it, with an owner, a sequence, and a target date. The roadmap is not a flat list. It is ordered, so that the enterprise knows what to do first, and it is grounded in the findings from the assessment, so that each item of work addresses a specific evidenced gap rather than a general aspiration to improve.
Sequencing the roadmap matters as much as its content. Some capabilities depend on others, so building them in the wrong order wastes effort. Visibility and accountability tend to come first, because assurance and runtime control depend on knowing what AI exists and who owns it. The roadmap respects these dependencies, so that foundational capabilities are built before the capabilities that rely on them. A roadmap that tries to build sophisticated assurance before basic visibility exists will stall, and the assessment is designed to prevent that.
The roadmap should be fundable, which means it should be expressed in terms a board can approve. Each item should carry an indication of the risk it reduces and the effort it requires, so that the board can make informed trade-offs. The flow below shows the path from the heat map through prioritisation to a sequenced roadmap and back to reassessment, which is the loop that turns a single assessment into a programme of continuous improvement rather than a report that is read once and shelved.
Heat map to roadmap to reassessment
The assessment turns a diagnosis into a sequenced, fundable plan, then verifies progress at the next cycle.
Current against target across domains, evidence backed.
Rank gaps by risk and by dependency.
Sequenced work with owners and dates.
Verify closure and rebaseline the heat map.
Prioritising by risk
Not every gap deserves equal attention, and the assessment prioritises by matching gap size to risk. A large gap in a low risk domain can wait, while a smaller gap in a high risk domain may be urgent. The principle is that the next investment should go where the exposure is greatest, which usually means the domains that manage autonomy, customer impact, sensitive data, or regulatory obligation. Prioritising this way ensures that scarce governance effort reduces the most risk per unit of work, which is what makes the programme defensible to a board.
Prioritisation also considers the cost of delay. Some gaps grow more dangerous the longer they remain open, particularly runtime control gaps in an enterprise that is rapidly expanding its use of agents. A gap that is tolerable today may be serious in six months as autonomy increases, so the assessment weighs not only current risk but the trajectory of risk. A gap that is small but growing may warrant earlier attention than a gap that is larger but stable, and the roadmap reflects this forward looking view.
The bars below show one illustrative view of gap size by domain, which helps explain how an enterprise might visualise the distance from target. It is a reference model, not measured data or an industry benchmark. No domain should be presumed weak: each enterprise must derive its own profile from evidence-based ratings, then combine gap size with the risk each domain manages to set roadmap priorities.
Illustrative capability gap by domain
The distance from target by domain. Combined with the risk each domain manages, this drives the ordering of the roadmap.
Evidence behind every score
The credibility of the assessment rests entirely on the evidence behind each score, so the assessment treats a score without evidence as a finding rather than a result. When a domain is rated managed, there should be records that show the capability operated consistently during the period. When those records do not exist, the honest rating is lower, no matter how well the capability is described. This discipline is what separates a real assessment from a self flattering questionnaire, and it is the single most important protection against a vanity score.
Requiring evidence also changes behaviour over time, which is a benefit beyond the assessment itself. When teams know that a high rating requires records, they build governance that produces records, which is exactly the governance an enterprise needs. The assessment thereby creates an incentive to make controls produce evidence as they operate, rather than to assemble evidence after the fact when an assessment or incident forces the search. Evidence produced in the moment is stronger and harder to dispute than evidence reconstructed later.
The strongest evidence is produced as a byproduct of control, particularly for the runtime domain. Records of actions that were approved, held, or blocked while AI was operating are far more convincing than a policy stating what should happen, because they show what actually did. This is the pattern of operational policy governance, and it is why enterprises that enforce policy at runtime find the assessment easier: the evidence already exists, produced by the controls themselves, and the assessment becomes a matter of examining it rather than hunting for it.
Benchmarking and comparability over time
A capability assessment is most powerful as a time series. A single result tells an enterprise where it stands, but a sequence tells it whether it is improving, and improvement over time is the real measure of a governance programme. To support this, the assessment holds the model, the method, the scale, and the targets stable across cycles, so that a change in the heat map reflects a change in governance rather than a change in how it was measured. Stability of method is what makes the trajectory trustworthy.
Comparability across business units is equally valuable, and it comes from the same discipline. When every unit is assessed against the same domains using the same method, the enterprise can aggregate results into a portfolio view and see which parts of the organisation are strong and which lag. This lets governance investment be directed across the enterprise rather than negotiated unit by unit, and it lets good practice in one unit be identified and spread. Comparability turns a set of local assessments into an enterprise capability.
External benchmarking should be approached with caution, because published maturity statistics are often inconsistent in method and difficult to verify. The assessment favours internal comparability over external benchmarking, because an enterprise can trust its own consistent method in a way it cannot trust a benchmark assembled from disparate sources. Where external comparison is useful, it should be treated as directional context rather than a target, and the enterprise should not adjust its own honest ratings to match an external figure whose basis it cannot see.
Presenting to the board and common mistakes
Presenting the assessment to the board is where its value is realised or lost. The board needs the heat map, a small number of severe findings with their remediation plans, and the trajectory against previous cycles. It does not need the full evidence file, but it should be able to reach into any rating to see the evidence behind it. Presentation that is only a score invites false comfort, and presentation that is only a list of findings overwhelms, so the assessment pairs the heat map with prioritised findings to give the board both shape and specifics.
The most common mistake is the vanity assessment, where scores are inflated to present a comfortable picture. This fails the enterprise twice, first by hiding real exposure and second by destroying the trajectory, because an inflated baseline cannot show honest improvement. The assessment guards against this by requiring evidence and independence, but the discipline has to be held against the real pressure to look good. An assessment that never finds a serious gap in a growing AI estate should be treated with suspicion, not relief.
A second mistake is to produce the assessment and stop. An assessment that is not connected to remediation, funding, and reassessment changes nothing, and the effort of producing it is wasted. The assessment is designed to feed the roadmap, the risk register, and board reporting, so that its findings drive change and its next cycle verifies it. An enterprise that assesses without acting has bought a diagnosis and declined the treatment, which is the least useful thing an assessment can be.
Conclusion: the Helixar perspective
The Helixar research perspective is that a capability score should be defensible under challenge, and defensibility comes from evidence. When each rating is backed by operational records of what controls actually did, the assessment survives scrutiny from a board, a regulator, or an external assurance provider. This is where operational policy governance and a strong evidence framework turn a self assessment into an audit ready record, because the evidence exists as a byproduct of enforcement rather than as a special collection exercise.
The assessment also keeps an enterprise honest about the domains it has built least. Runtime control and assurance consistently show the largest gaps, not because enterprises are careless but because these capabilities require more than a document to exist. Naming those gaps clearly, and sequencing the work to close them, is what turns an uncomfortable heat map into a credible programme. The assessment is at its best when it tells an enterprise something it did not want to hear and gives it a way to act on it.
Read alongside the capability model and the assessment methodology, this report completes the loop from what governance an enterprise needs, to how well it operates, to where to invest next. Together they let an enterprise measure and improve its AI governance with the same discipline it applies to its finances or its security. For the whole picture these reports support, the Enterprise AI Governance Framework is the anchor.
Enterprise checklist
- Apply the capability model and methodology, not an ad hoc questionnaire.
- Score current state on evidence, separating design from operating effectiveness.
- Set target state by risk, and keep targets stable across cycles.
- Produce a heat map, not a single vanity score, and report gaps honestly.
- Prioritise the largest gaps that meet the highest and fastest growing risk.
- Support every score with retained evidence, treating absence as a finding.
- Turn the heat map into a sequenced, fundable roadmap and reassess on a cycle.
Frequently asked questions
What does a capability assessment produce?
How do we avoid a vanity score?
Should every domain reach the top capability level?
Why favour a heat map over a single score?
How does the assessment stay comparable over time?
Method and source use
This report is a Helixar synthesis of the cited public standards and guidance. Named sources are linked at first mention and listed below. Unless a cited source is identified, maturity levels, diagrams, allocations, scores, and operating models are illustrative Helixar reference models, not survey findings or legal requirements. Organisations should verify current obligations with the authoritative source and qualified advisers.
References
- NIST AI Risk Management Framework (AI RMF 1.0)
- NIST AI RMF Playbook
- ISO/IEC 42001:2023, Artificial intelligence management system
- ISO/IEC 38507:2022, Governance implications of the use of AI by organizations
- ISACA COBIT
- AICPA SOC Suite of Services
- Helixar research: Enterprise AI Governance Capability Model
- Helixar research: Enterprise AI Governance Assessment Methodology
- Helixar research: AI Governance Maturity Model