AI agent governance maturity modelAI agent governanceEnterprise AI agentsAI agent controlsAI agent fleet managementAI agent context governance

AI Agent Governance Maturity Model

An AI agent governance maturity model measures how consistently an organization inventories, controls, monitors, audits, and improves agent work across its fleet.

Abe Wheeler
Alignbase wordmark on a deep blue background.
Alignbase wordmark on a deep blue background.

An AI agent governance maturity model measures how consistently an organization can govern agent work across teams, workflows, and time.

The model looks beyond whether a policy exists. It asks whether the organization knows which agents are running, assigns accountable owners, limits authority, governs context and data, tests controls, monitors production behavior, preserves evidence, responds to incidents, and improves those capabilities across the fleet.

Maturity is an organizational property. A single well-controlled agent does not make the fleet mature, and a high average score should not hide an unmanaged production workflow with broad access.

TL;DR

Use this AI agent governance maturity model in six steps:

  1. Define the fleet, business units, environments, and risk tiers in scope.
  2. Gather current operating evidence instead of scoring policy statements alone.
  3. Score eight governance dimensions from 0 to 4.
  4. Run the foundation diagnostic to catch a score that ignores a missing prerequisite.
  5. Record variation by team, workflow, and risk tier instead of forcing one flattering number.
  6. Fund the next capabilities that remove the largest risk and evidence gaps.

The five maturity levels are:

  • Level 0: Unmanaged
  • Level 1: Documented
  • Level 2: Repeatable
  • Level 3: Measured
  • Level 4: Continuously governed

The eight dimensions are inventory and ownership; purpose, risk, and autonomy; identity and authority; context, Skills, and Memory; testing and change control; runtime monitoring and response; audit and assurance; and fleet governance and improvement.

Model version: 1.0, published August 30, 2026. Scorecards should record this model version separately from the article’s modification date.

What Is an AI Agent Governance Maturity Model?

An AI agent governance maturity model is a staged way to assess the organization’s ability to direct, control, and account for agents. It turns a broad governance goal into observable capabilities and evidence.

The model answers four questions:

  1. What governance capability exists now?
  2. Is that capability used consistently across the assessed scope?
  3. What evidence shows that it works during real agent operation?
  4. Which missing capability should the organization build next?

The model is not a certification, legal opinion, or claim that every agent is ready for production. It is a management tool. Its value comes from exposing variation and setting an improvement sequence that owners can fund and verify.

The Cloud Security Alliance’s draft Agentic AI Governance Maturity Model also uses a five-level progression and assesses capabilities such as identity governance, runtime controls, tool management, human oversight, incident response, compliance, and workforce capability. It provides a useful reference structure, while each organization still needs evidence thresholds that match its agents, authority, data, obligations, and risk tolerance.

Why Agent Governance Needs Its Own Maturity Model

General AI governance still applies, but agents add control surfaces that a model inventory or review committee may not cover.

An agent can:

  • Receive changing instructions, Skills, Memory, retrieved data, and tool output
  • Plan and execute several steps without a person approving each one
  • Use a delegated user identity or its own service identity
  • Call tools that read, write, send, publish, spend, or execute code
  • Create schedules, retries, queued work, and persistent state
  • Delegate work to other agents and influence shared Memory
  • Produce an external effect that differs from its final text response

Those capabilities create a governance question at every step: who authorized this agent to do this work, under which current rules, with which evidence, and who can stop it?

IMDA first published its Model AI Governance Framework for Agentic AI in January 2026, then updated it to version 1.5 in May 2026. The framework groups practical guidance around bounding risk and authority, meaningful human accountability, technical controls across the lifecycle, and informed use. Version 1.5 also adds practices for multi-agent systems, third-party agents, and automation bias. A maturity assessment checks whether those kinds of measures are isolated intentions or repeatable operating capabilities.

Maturity, Readiness, Risk, and Assurance Answer Different Questions

These practices support each other, but they do not use the same subject or decision.

Practice Unit of review Main question Typical output
Governance maturity Organization, business unit, or fleet How repeatable and evidenced are our governance capabilities? Dimension profile, maturity level, gaps, and improvement plan
Readiness assessment One workflow and operating stage May this workflow enter its next stage? Scores, blockers, decision, and owned actions
Risk management Risk scenario, system, or portfolio What can go wrong, and how will we treat it? Risk register, treatment, monitoring, and acceptance
Security assessment One system or release scope Do required security controls work for this scope? Procedures, findings, decision, and retest record
Assurance Defined claim and decision Does the evidence justify confidence in this claim? Assurance case, limits, conclusion, and review triggers

A mature organization can still reject an unready workflow. A less mature organization may run one tightly controlled pilot, but it should not treat that success as proof that the same controls will appear across every team.

Define the Assessment Scope Before Scoring

Start by naming the organizational boundary. It may be the whole company, one business unit, one country, one product group, or one agent platform.

Record:

  • Business units, teams, and legal entities in scope
  • Production, pilot, development, and personal-use environments
  • Agent types, runtimes, integrations, and deployment models
  • User groups, affected people, and external parties
  • Data classes, tools, target systems, and external actions
  • Risk and autonomy tiers
  • Assessment period and evidence cutoff
  • Owners, reviewers, and decision-makers
  • Known exclusions and unavailable evidence

Do not hide scope gaps. If the central platform team has strong controls but business units deploy agents through unsupervised automation accounts, report both states. The lower score may apply only to that unit, but it is still part of the enterprise risk picture.

Segmenting by risk tier also matters because the assessment profile, evidence depth, and required controls should rise with authority and possible harm. Report maturity separately for materially different tiers. No maturity level authorizes an agent to operate; each workflow still needs its own readiness, risk, security, and release decisions.

Use Evidence, Not Confidence, to Assign Levels

Score what operated during the assessment period. Policies and architecture diagrams are evidence of intent. They do not prove consistent use or control effectiveness.

Use four evidence classes:

Design evidence

Policies, standards, ownership matrices, architectures, data-flow maps, role definitions, risk taxonomies, control requirements, and lifecycle procedures show what the organization expects.

Implementation evidence

Registry records, identity configuration, authorization policy, route configuration, tool allowlists, approval rules, test suites, deployment gates, log schemas, and stop controls show that the design has been built.

Operating evidence

Run records, authorization decisions, approval use, context delivery records, Skill reads, monitoring alerts, access reviews, lifecycle events, incidents, response exercises, and verified external outcomes show that the capability ran.

Effectiveness evidence

Coverage, exceptions, false-positive and false-negative analysis, control-failure rates, mean time to detect, mean time to revoke, reconstruction success, test escape rates, and remediation trends show whether the capability produces its required result.

The NIST AI Risk Management Framework organizes work around Govern, Map, Measure, and Manage. NIST also states that its Playbook is not a one-size-fits-all checklist. Use that principle here: adapt the evidence and thresholds to the scope, then preserve enough structure to compare the same capability over time.

Assign cumulative dimension scores

The fixed evidence crosswalk below is the scoring contract. The long evidence lists explain the controls and records referenced by a crosswalk cell; they are not separate scored criteria. Mark each requirement within a crosswalk cell as met, not met, or not applicable for the defined scope. A cell is met only when every applicable requirement is met and every not-applicable decision is approved. An approved not-applicable requirement counts as satisfied for cumulative scoring only when the defined scope makes that capability irrelevant. Each decision needs a recorded reason, evidence, and reviewer. A missing capability cannot become not applicable only because the organization has not built it.

The detailed lists define the evidence terms used in the crosswalk. For example, Level 2 approval integrity includes canonical rendering, digest binding, approver eligibility, separation of duties when required, atomic use, expiry, replay resistance, and idempotency. Level 2 dangerous-test safety includes written authorization and an isolated environment. Level 2 context evidence distinguishes server assembly, trusted session insertion, Skill package availability, and required invocation. These definitions are fixed parts of the named crosswalk cell, not criteria an assessor may move to another level.

For each dimension, assign the highest score whose requirements and every lower level are met:

  • Score 0: A material part of the dimension is unmanaged, unknown, or unsupported by reliable evidence.
  • Score 1: Applicable ownership, scope, policy, and baseline criteria are documented and implemented across the assessed scope.
  • Score 2: All Score 1 criteria are met, required controls and processes are repeatable, and current implementation and operating evidence covers the assessed scope.
  • Score 3: All Score 2 criteria are met, effectiveness measures are current, thresholds drive recorded decisions, and material exceptions are detected and either corrected or controlled by an approved equivalent compensating control.
  • Score 4: All Score 3 criteria are met, measured improvement is demonstrated, approved automation stays within its security limits, and the organization can maintain coverage as the fleet changes.

An open exception to a required criterion keeps the dimension below that level unless the assessment profile explicitly permits an exception and a time-bound compensating control produces equivalent evidence. When those conditions are met, count the requirement as satisfied only for the approved exception period. A requirement that law, contract, risk policy, or the approved profile marks non-waivable cannot receive an exception. Without approved equivalent compensation, record the lower score until the required result is restored.

Score Eight AI Agent Governance Maturity Dimensions

Score each dimension separately from 0 to 4. Write the reason, evidence, exceptions, owner, and next action next to the number.

1. Inventory, ownership, and lifecycle

Assess whether the organization can discover and track agents from proposal through retirement.

Evidence should show:

  • A maintained registry for production, pilot, scheduled, local, and delegated agents
  • Stable agent identity and purpose
  • Direct active human Owners
  • Business, technical, risk, data, and incident duties
  • Environment, model, tools, data, autonomy, and target systems
  • Lifecycle state, review date, and retirement status
  • Discovery of unmanaged or stale deployments
  • Suspension and retirement that stop new, queued, scheduled, in-flight, and delegated work as required

A spreadsheet updated before an annual audit may support Level 1. Automated discovery reconciled with an owned system of record and tested retirement evidence may support Level 3 or 4.

2. Purpose, risk, and autonomy

Assess whether each agent has an approved purpose, operating envelope, risk tier, and authority boundary.

Look for:

  • Intended and prohibited uses
  • Users, affected people, and business outcomes
  • Data, tool, cost, duration, volume, and external-effect limits
  • Defined autonomy levels
  • Risk assessment tied to actual authority and action chains
  • Required human or independent checkpoints
  • Accepted remaining risk with a named human decision-maker
  • Reassessment after material changes or incidents

The score should reflect enforcement and current review, not how clear the risk policy sounds.

3. Identity, authorization, tools, and data

Assess whether agents act through distinct, revocable identities and whether trusted boundaries check authority before protected operations.

Evidence should cover:

  • Agent, initiating user, service, and peer-agent identity
  • Authenticated workload identity and server-derived principal and tenant context from validated credentials, not caller-supplied identity fields
  • Scoped, short-lived credentials where practical
  • Action-time authorization against current identity, tenant, purpose, target, parameters, environment, and approval
  • Authorization at every delegation hop, with effective authority limited to the intersection of the initiating principal, sending agent, receiving agent, task grant, target policy, and current approval
  • Audience-bound delegation credentials, independent receiver eligibility checks, and rejection of forwarded credentials or grants that exceed the delegated task
  • Least-privilege tool and data access
  • Tool source, owner, version, operation allowlist, schema, and network destination
  • Trusted rendering of the exact canonical action, target, and material parameters; approval bound to their digest; authenticated approver eligibility and separation of duties where policy requires it
  • Approval atomic use, expiry, replay resistance, and idempotency
  • Tenant and data-class isolation
  • Fast credential and authorization-grant revocation
  • Denied-path tests with no prohibited side effect
  • Impersonation, audience-confusion, altered-claim, and cross-tenant authorization tests

Prompt instructions can explain a limit to the agent, but they do not replace authorization at the protected API, tool, data, or execution boundary.

4. Context, Skills, Memory, and delivery

Assess whether the organization manages agent inputs as governed context artifacts rather than scattered prompt fragments.

Evidence should show:

  • Named Owners and scoped repository access roles
  • Review and controlled release for stable instructions and Skill packages
  • Live, versioned, and audited working Memory with controlled writers and readers
  • Separate repository access from delivery assignments or routes
  • Delivery authorization, target eligibility, tenant limits, and change history
  • Expected released instruction content, current Memory content, and released Skill metadata and package digest in the server bundle
  • Delivery-exclusion evidence showing artifacts barred by delivery policy, tenant scope, target eligibility, or explicit non-routing are absent from the automatic bundle
  • Denied-read evidence showing repository access rules independently prevent unauthorized discovery or reads from injecting content into a session
  • Trusted request- and session-bound post-insertion evidence for instruction and Memory content, or equivalent downstream session evidence; a host that provides neither creates an evidence gap
  • Package fetch or local sync evidence that the available Skill package matches the advertised digest
  • Run-bound invocation evidence tied to that package digest when the workflow requires the Skill to be used, with availability and invocation reported as separate results
  • Point-in-time reconstruction of inputs for material runs
  • Detection of stale, conflicting, excessive, missing, or cross-tenant context

At higher maturity, teams also measure bundle size, duplication, freshness, required-content coverage, delivery failure, and whether changes improve evaluated outcomes.

In Alignbase, these general controls map to Knowledge and Skills with review and publication, live versioned Memory, repository permissions, Always routes, and read_skill or local sync for Skill package availability.

5. Testing, release, and change control

Assess whether agent changes pass repeatable tests and release gates matched to risk.

Look for:

  • Normal, edge, adversarial, dependency-failure, recovery, and multi-agent cases
  • Tests for prompt injection, tool misuse, Memory poisoning, approval bypass, data leakage, loops, retries, and unsafe delegation
  • Exact model, context, Skill, Memory policy, tool, permission, and runtime versions under test
  • Independent review for high-impact workflows
  • Hard blockers that override average test scores
  • Written authorization and isolated environments for dangerous tests
  • Staged rollout, monitoring, rollback, and verified cleanup
  • Material-change triggers and retest rules
  • Evidence that the deployed release matches the assessed release

The OWASP State of Agentic AI Security and Governance 2.01 collects current security and governance guidance for autonomous systems. Use threat guidance to shape tests, but score the governance capability on whether tests are required, safe, repeatable, reviewed, and tied to release decisions.

6. Runtime monitoring, incident response, and recovery

Assess whether operators can detect a material deviation, contain it, preserve evidence, restore a safe state, and learn from the event.

Evidence should include:

  • Monitoring of tool calls, external effects, authority use, context changes, policy decisions, errors, cost, duration, and delegation
  • Alerts tied to owners and response targets
  • Detection of missing evidence and silent control failure
  • Tested stop, suspend, quarantine, credential revocation, scoped context access-grant revocation, and delivery-assignment removal paths, with each control changed and verified independently
  • Independent confirmation that side effects and delegated work stopped
  • Agent-specific incident classification and runbooks
  • Evidence preservation outside the acting agent’s authority
  • Recovery, reconciliation, compensation, notification, and review
  • Feedback from incidents and near misses into tests, policies, Skills, and controls

General infrastructure monitoring may help, but higher maturity requires agent-aware signals connected to decisions and external outcomes.

7. Audit, accountability, and assurance

Assess whether another reviewer can reconstruct important agent work and determine who approved, executed, changed, and owned it.

Evidence should connect:

  • Initiating principal and agent identity
  • Workflow, run, session, and external record IDs
  • Context, Skill, Memory, model, tool, and policy versions
  • Route, permission, authorization, and approval decisions
  • Tool requests, results, retries, and verified external effects
  • Human review, exceptions, risk acceptance, and release decisions
  • Changes, lifecycle events, incidents, and remediation
  • Retention, integrity, access, and export controls for evidence

Server-side assembly records prove what the server prepared. They do not alone prove what a client inserted into a model session or which Skill package was available. Mature audit design states what each record proves and reconciles evidence across the full path.

8. Fleet governance, measurement, and improvement

Assess whether governance works across the fleet instead of through isolated projects.

Look for:

  • Common policies, risk tiers, evidence contracts, and minimum controls
  • Local ownership within centrally defined boundaries
  • Fleet views by agent, owner, business unit, environment, data class, and authority
  • Coverage and effectiveness measures for every material control
  • Exceptions with owners, expiry, and compensating controls
  • Cross-team incident and evaluation learning
  • Budget and staff for governance capability
  • Training for builders, owners, reviewers, and responders
  • Versioned changes to the governance model itself
  • Evidence that improvement work reduced a measured gap

Level 4 does not require governance to rewrite itself without human control. Automated tuning may tighten controls or adjust non-security settings inside approved limits, but it must not expand authority, broaden tenant or data scope, reduce required approval, weaken mandatory tests, or remove a release blocker without an accountable human decision. Every automated change should remain reversible, produce a decision record, fail closed when its policy or evidence is unavailable, and receive independent verification that the resulting control state matches the approved bounds.

Minimum Evidence by Dimension and Level

Use these level tables to make the dimension scores reproducible. Each row states the dimension-specific requirement for that level. A score also requires the global definition for that score, every lower-level row in the same dimension, and the fixed evidence definitions above for terms used in that row. For example, a Level 1 row that says a control is recorded still requires it to be implemented across the assessed scope, and a Level 3 row that says results are reviewed still requires thresholds to drive recorded decisions and material exceptions to be corrected or controlled by an approved equivalent compensating control.

Inventory and lifecycle

Level Minimum evidence
1 Current observation sources are reconciled with the registry, and all in-scope agents have active human Owners, purpose, authority, and lifecycle state.
2 Discovery, provisioning, review, suspension, and retirement follow one tested process across the scope.
3 Coverage, stale-agent, revocation, and retirement results are measured and exceptions trigger action.
4 Near-current discovery and lifecycle checks keep pace with fleet change, and measured fixes reduce gaps.

Purpose, risk, and autonomy

Level Minimum evidence
1 Risk tiers, intended and prohibited uses, authority bounds, and decision owners are recorded.
2 The same tiering, control profile, checkpoint, and reassessment process applies to comparable agents.
3 Exceptions, operating-envelope violations, outcomes, and accepted risk are measured and reviewed.
4 Evidence improves profiles and oversight, while material authority or risk changes still require accountable human approval.

Identity, authority, tools, and data

Level Minimum evidence
1 Authenticated distinct and revocable identities, enforced allowed-tool and data scopes, approval rules, and tenant boundaries are documented and implemented.
2 Trusted boundaries derive principal and tenant context from validated credentials; action-time authorization, audience-bound delegation, canonical approval integrity, isolation, impersonation, and denied paths pass current tests across the scope.
3 Authorization, bypass, privilege, approval, and isolation results are measured and reconciled with external effects.
4 Measured improvements reduce authorization and isolation gaps; automated conformance may tighten access, while any authority expansion or control reduction requires human approval and reassessment.

Context, Skills, Memory, and delivery

Level Minimum evidence
1 Context artifacts, Owners, access roles, delivery assignments, versions, release rules, and Memory lifecycle are recorded.
2 Required server assembly, trusted session evidence, Skill package fetch or sync, required invocation, delivery-exclusion, denied-read, and reconstruction checks pass.
3 Coverage, freshness, conflict, delivery, invocation, reconstruction, and evidence-gap rates drive action.
4 Automated conformance may quarantine stale or unauthorized inputs, and measured changes improve outcomes without weakening human control.

Testing, release, and change

Level Minimum evidence
1 Risk-based test, review, blocker, change, and release requirements are written and assigned.
2 Applicable normal, adversarial, failure, recovery, multi-agent, and dangerous-test safety requirements pass, with enforced blockers, material-change retesting, staged rollout, rollback, cleanup, and deployed-versus-assessed matching.
3 Coverage, escape, recurrence, rollback, and deployed-versus-assessed match rates are measured and reviewed.
4 Measured improvements reduce test and release gaps; new incidents and attack cases enter fleet tests, while automated changes may add tests but cannot remove required tests without human approval.

Monitoring, response, and recovery

Level Minimum evidence
1 Required signals, alerts, owners, runbooks, stop paths, and evidence duties are documented and implemented.
2 Detection, suspension, credentials, context access grants, delivery assignments, delegated work, recovery, and cleanup pass exercises across the scope.
3 Detection, containment, revocation, recovery, recurrence, and missed-signal measures drive recorded action.
4 Near-current signals detect drift early, and verified improvements reduce response and recovery gaps.

Audit, accountability, and assurance

Level Minimum evidence
1 Evidence fields, stable IDs, owners, retention, access, integrity, and review duties are defined.
2 Material samples reconstruct identity, authority, context, Skills, approvals, tool use, and external effects end to end.
3 Completeness, reconstruction, evidence-gap, exception, and remediation measures drive assurance decisions.
4 Measured improvements reduce evidence gaps, while portable evidence and independent review keep assurance current as platforms, controls, and the fleet change.

Fleet governance and improvement

Level Minimum evidence
1 Common policy, control profiles, ownership, training, exception rules, and governance funding exist.
2 Comparable agents follow common processes, and coverage reports expose adoption gaps across teams.
3 Fleet control effectiveness, exception age, incident trends, and remediation results guide priorities and funding.
4 Versioned governance changes produce measured improvement, with bounded automation and human ownership of material decisions.

The Five Maturity Levels

Use these level descriptions as interpretive summaries, not as separate scoring criteria or labels for organizational status. Only the global score definitions and dimension crosswalk above determine a score.

Level 0: Unmanaged

Agents appear through team experiments, personal accounts, vendor features, and automation scripts without a complete inventory or shared control model.

Typical signs include:

  • Shared credentials or activity attributed only to a human account
  • Unclear ownership and purpose
  • Prompt-only rules
  • Scattered context and Skills with no version history
  • Broad tool access
  • Tests focused on happy-path output
  • Logs that cannot reconstruct authority, inputs, and effects
  • Incident handling after visible harm

The priority at Level 0 is discovery and containment. Find the agents, assign owners, identify high-impact authority, and stop unmanaged production access before building a large framework.

Level 1: Documented

The organization has an inventory, named owners, initial risk tiers, written policies, and basic review points. Coverage may still depend on team discipline and manual checks.

Typical evidence includes:

  • Current observation sources reconciled with a reviewed registry for all in-scope agents
  • Approved purpose and owner fields
  • Basic data and tool boundaries
  • Written context, testing, approval, and incident rules
  • Manual release checklists
  • Central log collection
  • Scheduled access and lifecycle reviews

Level 1 creates visibility and responsibility. It does not yet prove that every deployment follows the same process or that important controls cannot be bypassed.

Level 2: Repeatable

Teams use a common lifecycle, control baseline, evidence contract, and release process. Important controls run at trusted boundaries, and exceptions follow a defined path.

Typical evidence includes:

  • Registry and identity provisioning connected to deployment
  • Risk-based control and test requirements
  • Action-time authorization and enforced approval gates
  • Governed context, Skill, Memory, and route lifecycles
  • Standard test suites and hard release blockers
  • Point-in-time run and configuration records
  • Tested suspension, incident, and retirement procedures
  • Coverage reports that find missing adoption

At Level 2, the same kind of agent should receive the same minimum governance regardless of which team built it.

Level 3: Measured

The organization measures whether controls work in operation and uses that evidence to decide where to intervene.

Typical evidence includes:

  • Current fleet coverage by control and risk tier
  • Control success, failure, bypass, and evidence-gap rates
  • Mean time to detect, contain, revoke, reconstruct, and remediate
  • Production outcome reconciliation
  • Drift and change-impact measures
  • Independent sampling and assessment
  • Exception age and recurrence trends
  • Funding decisions tied to measured risk and control performance

Level 3 moves governance from process adoption to operating effectiveness. A dashboard alone does not qualify. Measures need definitions, owners, data-quality checks, thresholds, and decisions that follow when a threshold is missed.

Level 4: Continuously Governed

Governance improves through current fleet evidence, incidents, evaluations, and business change. Approved automation can adjust bounded controls, while material policy and risk decisions remain accountable to people.

Typical evidence includes:

  • Near-current discovery and governance coverage across environments
  • Policy-as-code and automated conformance where suitable
  • Automated security changes only tighten controls; any reduction in oversight, wider limit, or weaker test requirement receives accountable human approval and reassessment
  • Early detection of control and behavior drift
  • Cross-fleet replay of relevant incident and attack cases
  • Measured improvement after control changes
  • Portable evidence and policy across agent platforms
  • Versioned governance methods with independent review

Level 4 is not a claim of perfect control. It means the organization can detect new gaps, change the right controls, verify the result, and preserve accountability as the fleet changes.

Assign the Overall Maturity Level Without Hiding Gaps

An average score is useful for trend reporting, but it should not determine the operating decision by itself.

Use three outputs:

  1. Dimension scores from 0 to 4
  2. A maturity distribution by business unit, environment, and risk tier
  3. An overall level equal to the lowest dimension score

Set the overall level to the lowest of the eight cumulative dimension scores. Report the average or median only as a secondary trend measure. This rule prevents measurement, automation, or one strong control domain from raising the overall level above an unresolved lower-level gap.

Use the foundation diagnostic below to check the calculation when the assessed scope lacks any of these foundations:

  • A current enough inventory to define the fleet
  • Active human ownership for production agents
  • Approved purpose, risk, and authority boundaries
  • Distinct and revocable production identity
  • Action-time authorization for protected operations
  • Required controls and safe-stop capability for high-impact work
  • Material-change and incident processes
  • Reconstructable evidence for consequential actions

The diagnostic is not a second scoring algorithm. Its expected maximums restate where the fixed crosswalk should place the affected dimension. If a calculated level exceeds one of these values, revisit the dimension evidence because it was scored incorrectly.

Missing foundation Maximum level Reason
Current enough inventory or active human ownership 0 The organization cannot define or account for the fleet it claims to govern.
Approved purpose, risk tier, or authority boundary 0 The scope lacks the Level 1 basis for selecting and judging controls.
Authenticated distinct and revocable identity, enforced tool and data scope, or enforced tenant boundary 0 The scope lacks the Level 1 basis for attribution, revocation, or isolation.
Action-time authorization, tenant-isolation test evidence, or denied-path test evidence 1 Baseline enforcement exists, but repeatable protection has not been proved.
Per-hop delegated-authority attenuation or repeatable approval-integrity evidence 1 Baseline limits exist, but repeatable delegation or approval integrity has not been proved.
Required stop, suspension, credential revocation, context access-grant revocation, delivery-assignment removal, or incident capability 0 The scope lacks the Level 1 capability to contain an affected agent.
Required containment or recovery exercise does not pass 1 The capability exists, but repeatable operation has not been proved.
Repeatable lifecycle, material-change, testing, and release process 1 Governance still depends on local judgment rather than a shared operating process.
Reconstructable evidence for consequential actions, required context and Skill delivery, or required Skill invocation 1 The organization cannot prove repeatable operation across the full action path.

For example, an organization with strong testing and monitoring but no reliable agent inventory remains at Level 0 for that scope because the inventory dimension is 0. A business unit may have a higher profile for a bounded, fully inventoried subset, but the report must state that narrower scope.

A compact scorecard can look like this:

Dimension Score Evidence reviewed Material gap Owner Target date
Inventory and lifecycle 0 Registry export, discovery scan, retirement drill Local scheduled agents outside inventory Platform 2026-10-15
Risk and autonomy 0 Policy, sample risk decisions Agents lack approved risk tiers and authority boundaries Risk 2026-09-30
Identity and authority 0 Identity records, authorization tests Two legacy shared accounts Security 2026-09-20
Context and routing 1 Context artifact history, route records, delivery samples No trusted session evidence for one host AI platform 2026-10-31
Testing and change 1 Release gates, test reports, rollback exercise Multi-agent cases incomplete Engineering 2026-10-10
Monitoring and response 1 Logs, alerts, incident drill Stop drill misses scheduled work Operations 2026-09-15
Audit and assurance 1 Run samples, approval records External effects do not reconcile Assurance 2026-10-01
Fleet improvement 1 Common policy, quarterly review No common process or adoption coverage report Governance council 2026-11-01

The sample’s overall level is 0 because the lowest dimension score is 0. The foundation diagnostic confirms that inventory, risk-tier, and identity gaps belong at Level 0, while the failed stop exercise and evidence gaps belong at Level 1. The stronger controls elsewhere still matter, but they cannot support a higher fleet claim while part of that fleet remains outside the inventory.

The table is a profile, not a report card for individual teams. Use it to fund and sequence work. Keep sensitive findings and system detail in access-controlled records, then link those records from the scorecard.

Build the Next Level in Dependency Order

Maturity programs fail when they start with advanced analytics before fixing inventory, ownership, and enforcement.

Use this dependency order:

  1. Discover agents and bound high-impact authority.
  2. Assign human ownership and common risk tiers.
  3. Establish identity, authorization, context, tool, data, and approval controls.
  4. Create repeatable lifecycle, testing, release, incident, and evidence processes.
  5. Measure coverage and effectiveness.
  6. Improve controls from operating evidence.

First 30 days: establish the baseline

  • Define scope and risk tiers.
  • Reconcile an agent inventory from deployment, identity, API, tool, expense, and network sources.
  • Assign active human Owners.
  • Mark agents with production access, sensitive data, write authority, code execution, public output, or delegation.
  • Contain unknown or unowned high-impact agents.
  • Score the eight dimensions with evidence and record exclusions.

Days 31 to 60: make controls repeatable

  • Publish minimum controls by risk and autonomy tier.
  • Connect identity, registry, and lifecycle events.
  • Put authorization and approval checks at trusted boundaries.
  • Define governed context, Skill, Memory, and route processes.
  • Add test, release, change, stop, and incident requirements.
  • Define the evidence contract for consequential runs.

Days 61 to 90: measure operation

  • Measure control and evidence coverage across the fleet.
  • Sample runs and reconcile requested work, authority, context, tool calls, and external effects.
  • Exercise suspension, credential revocation, scoped context access-grant revocation, delivery-assignment removal, response, and recovery as independent controls.
  • Track exceptions and remediation age.
  • Re-score the dimensions and verify whether the chosen controls changed the evidence.

The next assessment should use the same definitions unless a versioned method change is necessary. Otherwise, a better score may reflect easier criteria rather than better governance.

Track Measures That Change Decisions

Useful measures connect a governance requirement to an action.

Examples include:

  • Percentage of observed agents registered and owned
  • Percentage of high-impact agents with distinct identity and current risk review
  • Percentage of protected actions covered by action-time authorization
  • Required-context assembly and trusted insertion evidence coverage
  • Skill package digest match and required invocation coverage
  • Material runs with complete decision and external-effect evidence
  • Release changes with current tests and approved rollout records
  • Mean time to suspend, revoke credentials, remove context access, remove delivery assignments, and verify stop
  • Control failures, bypass attempts, evidence gaps, and recurring exceptions
  • Mean time to remediate findings by severity
  • Percentage of retired agents with verified credential, schedule, route, and data cleanup

Avoid measures that reward volume without testing effectiveness. Counting policies, tests, log events, or training sessions can show activity, but it does not prove that agents stayed inside approved authority.

Every measure should have an owner, source, calculation, refresh period, target, alert threshold, and required decision. If nobody acts when a target is missed, the measure is reporting rather than control.

Keep the Maturity Model Current

Review the model at a fixed cadence and after material changes to the fleet, business, law, or threat environment.

Version changes when you:

  • Add a new agent type, host, model, tool protocol, or execution environment
  • Expand agents to a new business unit, country, data class, or affected population
  • Change autonomy, delegation, or human oversight patterns
  • Learn from an incident, near miss, control failure, or independent assessment
  • Adopt a new requirement, standard, or contractual duty
  • Find that a measure does not predict the outcome it was meant to control

Keep the old criteria and scores with the new version. Historical comparison needs to distinguish an actual capability change from a scoring-method change.

Version history

  • Version 1.0, August 30, 2026: Initial five-level model, eight-dimension evidence crosswalk, foundation diagnostic, and example scorecard.

How Alignbase Supports Governance Maturity

Alignbase is an AI context control plane and a context, Skills, and Memory repository. It supports the context side of the maturity model by giving teams governed Resources, scoped roles, independent routes, version history, and delivery records.

That progression can move from scattered prompts to:

  • Direct human ownership and Group-based access
  • Reviewed and published Knowledge and Skills
  • Live, versioned, and audited Memory
  • Always routes managed separately from repository permissions
  • Current context assembled for each agent
  • Point-in-time records for which versions the server delivered
  • Fleet-wide context coverage and change evidence

Alignbase does not replace identity, authorization, tool enforcement, monitoring, legal review, or incident response. It gives those systems a governed context layer, which is necessary because agents cannot follow a current rule they never receive.

The Standard to Aim For

A mature governance program can answer two questions at the same time:

  1. What controls should every agent at this risk level have?
  2. What evidence shows those controls worked for this agent, release, and action?

The first question creates consistency across the fleet. The second prevents consistency from becoming a paper claim.

Use the maturity model to expose the distance between them, assign owners to that gap, and verify that the next round of work changed how agents operate.

Frequently Asked Questions

What is an AI agent governance maturity model?

An AI agent governance maturity model measures how consistently an organization inventories agents, assigns ownership, bounds authority, governs context and data, tests controls, monitors operation, preserves evidence, handles incidents, and improves those capabilities across its agent fleet.

What are the levels of AI agent governance maturity?

This model uses five levels: Level 0 Unmanaged, Level 1 Documented, Level 2 Repeatable, Level 3 Measured, and Level 4 Continuously governed. Each level requires working evidence across eight governance dimensions, not only written policy.

How do you assess AI agent governance maturity?

Define the fleet and operating scope, gather current evidence, score eight dimensions from 0 to 4 with the fixed evidence crosswalk, run a foundation diagnostic to catch inflated scores, record variation by business unit and risk tier, and choose a small set of owned improvements for the next review period.

Is an AI agent maturity assessment the same as an agent readiness assessment?

No. A maturity assessment measures repeatable organizational capability across many agents and over time. A readiness assessment decides whether one defined agent workflow has enough evidence and controls to enter its next operating stage.

What evidence proves AI agent governance maturity?

Useful evidence includes a current agent registry, direct human Owners, risk and autonomy decisions, identity and authorization records, context and Skill versions, Memory history, route records, test results, approval decisions, run traces, external outcome checks, incidents, lifecycle events, and measured control performance.

Can an organization average its scores into one maturity level?

Averages can summarize a profile, but they should not hide a foundational gap. Missing inventory, ownership, action-time authorization, required controls, stop capability, or reconstructable evidence should cap the overall level for the affected scope even when other dimensions score well.

How does context governance affect AI agent governance maturity?

Context governance matures from scattered prompts to owned and versioned context artifacts, controlled release for stable instructions and Skills, audited live working Memory, access-independent delivery management, point-in-time reconstruction, and measured delivery evidence across the fleet.