Framework · September 2026

The Agentic Trust Maturity Model

A five-level scale for how disciplined your organization is at detecting, classifying and controlling the AI agents now acting on your customers’ behalf.

01InitialMost today
02RepeatableMost today
03DefinedNear-term goal
04ManagedRare
05OptimizingRare

Free. No sign-up, no form

Arkose Labs / Product Marketing

The Agentic Trust Maturity Model

Arkose Labs’ five-level model for how companies detect, classify, and control AI agents shopping, browsing, and transacting on behalf of their customers.

Draft of 18 September 2026

Executive Summary

AI agents are now acting on behalf of your customers; they are logging in, opening accounts, browsing, transacting, and consuming your services. Most organizations cannot say how much of their traffic this represents. 94% of the 804 security and business leaders surveyed for the World Economic Forum’s Global Cybersecurity Outlook 2026 named AI the most significant driver of change in cybersecurity in the year ahead.

This demands a new strategy because agentic traffic invalidates the assumption every control in place today was designed around. Our controls are designed to detect a bot and block it. Our controls are not designed to assess its intent or differentiate between the ones who adhere to acceptable policies and who don’t. Agentic AI use is only bound to increase. Treat them as bots and you block your own customers. Treat them as automated actors helping your customers and you risk opening yourself to automated abuse at scale.

The consequence for strategy is that agent handling stops being a straightforward control you configure and becomes a capability you build, spanning detection, classification, enforcement, and business ownership. Capabilities have to be measured over time. Agent frameworks, agent behavior, and the identity standards governing them are all evolving at a rapid rate, which means a point-in-time assessment of your capabilities is stale before the next board meeting. What leadership needs is not a score. It is a repeatable way to ask the same questions each quarter and see whether the organization has actually moved.

This report introduces the Agentic Trust Maturity Model: a five-level scale, Initial, Repeatable, Defined, Managed, Optimizing, assessed across five operational dimensions: visibility and detection, classification, enforcement, accountability, and standards integration. The level names follow the maturity-model conventions in use; everything inside the levels is built specifically for agentic traffic.

The model is designed to do three things. First, it establishes where an organization stands today holistically across each maturity level. Second, it identifies which gap to close first, most organizations are comparatively strong on detection and weak on classification and ownership, and the sequence matters. And third, it gives leadership a fixed baseline to re-measure against, so progress can be reported and funded like any other operational program. A crosswalk to NIST CSF 2.0 is available in the appendix, so this runs alongside the frameworks your security and risk teams already report against rather than adding another.

Most organizations are just starting their maturity journey, and that is to be expected and not a failure. Agent trust capabilities and standards only started taking form in the last 12 months. This framework is designed to help companies address the two failures we see most often: legitimate customers and legitimate revenue blocked as false positives, and adversarial agents missed because they no longer look like bots.

The case for moving is not defensive. Organizations that can tell quickly and confidently which agents are acting for their customers and which are acting against them will be able to say yes to agentic traffic as a channel while competitors are still deciding whether to block it. Maturity here is not a compliance milestone. It is the precondition for treating agents as customers rather than as threats.

Why Agentic Traffic Breaks the Old Playbook

Consumers have already started delegating shopping tasks to AI. Visa reports 47% of U.S. shoppers already using AI for at least one shopping task, and Gartner projects 40% of enterprise applications will embed AI agents by the end of 2026 (Visa, 2026 agentic commerce forecast; Gartner via commercetools, Agentic Commerce Stats 2026). Every digital platform with a customer account is already receiving agentic traffic today, whether or not it has decided what to do about it.

Harvard Business Review named the underlying problem plainly in February 2026: companies will capture the agentic opportunity only if they can earn, and keep, customer trust. The analysis identified five concrete failure modes eroding it (Furman, Gürdeniz, Safari, and Ural):

None of the five is specific to retail. Each describes an agent acting for a customer on any platform holding an account, a balance, a policy, or a booking. Three are maturity problems, and this report returns to them at the close.

Delegation is running ahead of authorization. Only 14% of U.S. consumers say they trust AI to place orders on their behalf, and the figure rises only to 29% among Gen Z and 30% among millennials (MetaRouter, Agentic Commerce Trends and Statistics for 2026). A separate consumer survey found 55% of shoppers are not comfortable letting an AI agent complete a transaction on their behalf, with fraud named as the leading concern (The Wise Marketer, coverage of a 2026 Riskified consumer study).

Figure 1
47% already use AI for at least one shopping task 14% trust it to place an order unsupervised 33 points the opportunity, and the risk
The Trust Gap. Sources: Visa 2026 (adoption, any single shopping task); MetaRouter 2026 (trust). commercetools 2026 reports a narrower 38% for actively AI-assisted shopping.

The detection problem is just as unresolved. Traditional bot management was built on a binary: is this session a human in a browser, or a scripted bot? Agentic traffic collapses the binary. A modern AI shopping agent can run on a real device, from a real residential IP, inside a real browser session, acting with the account holder’s full authorization. It looks nothing like the bots bot management was built to catch, and yet it is not a human either.

Figure 2
THE OLD MODEL Human or Bot Real device, real IP, real browser, account-holder authorization: an agent passes as human. WHAT IS ACTUALLY ARRIVING Human Direct, unassisted session Disclosing agent Identifies itself. Partner or verified integration. Non-disclosing agent Legitimate, acting for a real customer. Unannounced. Adversarial agent Fraud swarm, ATO, scaled abuse. The hard problem is the middle two. Legacy tooling treats both as bots, so legitimate revenue is blocked and adversaries are missed at the same time.
Why the Old Binary Collapses. The hard problem is the middle two. Legacy tooling treats both as bots, so legitimate revenue is blocked and adversaries are missed at the same time. Almost all agentic volume sits in the middle box, since disclosure has no ecosystem support behind it yet.

The result is a visibility gap the security industry is now measuring directly: only 24.4% of organizations report full visibility into the AI agents operating in their environment, more than half of observed agents run with no security oversight or logging at all, and 88% of organizations had a confirmed or suspected AI agent security incident in the past year (Gravitee, State of AI Agent Security 2026 Report). The gap is as much comprehension as tooling. In banking, a third of the CEOs, board members, and chief risk officers surveyed for Bank Director’s 2026 Risk Survey said they do not understand agentic AI at all, which makes this a governance problem before it is a detection problem.

Figure 3
24.4% of organizations report full visibility into AI agents in their environment 88% had a confirmed or suspected agent security incident in the past year 1 in 3 of bank leaders do not understand agentic AI at all Gravitee, State of AI Agent Security 2026 (first two). Bank Director 2026 Risk Survey, sponsored by Baker Tilly (third).
The Visibility Gap. Gravitee, State of AI Agent Security 2026 (first two). Bank Director 2026 Risk Survey, sponsored by Baker Tilly (third).

The payments industry is responding with new infrastructure rather than waiting for the problem to resolve itself. Google’s Agent Payments Protocol (AP2) uses signed Intent, Cart, and Payment mandates to let an agent transact within an authorized scope; Mastercard’s Agent Pay and Verifiable Intent, and Visa’s Trusted Agent Protocol, add network-level verified agent identity on top of card rails. In May 2026, Google and Mastercard contributed AP2 and Verifiable Intent to the FIDO Alliance, which has since stood up dedicated working groups to turn them into an interoperable industry standard (FIDO Alliance; PYMNTS, 2026).

This is meaningful progress on identity. It does not, by itself, tell a company what policy to apply to the agentic traffic already hitting its site today, across every endpoint, from every AI framework, disclosed or not. This is a maturity problem, not only a technology problem, and it is the gap this report addresses.

What Is the Agentic Trust Maturity Model?

The Agentic Trust Maturity Model asks one question: how disciplined is your organization’s process for detecting, classifying, and responding to non-human traffic, specifically the AI agents that may or may not be acting on your customers’ behalf? It answers on a five-level scale, from ad hoc handling dependent on individuals to a process documented, quantitatively controlled, and continuously improved.

Five Levels at a Glance

Table 1
LEVEL NAME IN ONE SENTENCE 1 Initial Ad hoc Agentic traffic is invisible inside general bot noise; nobody owns the problem. 2 Repeatable Reactive Legacy bot rules are reused for agents, inconsistently, by whichever team bought the tool. 3 Defined Standardized A documented, cross-functional policy classifies agent populations consistently across the business. 4 Managed Quantified Agent trust runs as a measured program with real-time visibility and graduated enforcement. 5 Optimizing Adaptive Agent trust is a continuously improving, revenue-generating capability, not a security cost center.
The five levels, in one sentence each.

Each level is assessed across the same fixed set of dimensions, applied identically throughout.

Figure 4
1 Visibility and detection Can you tell which sessions are agentic at all? 2 Classification Legitimate or adversarial, and trying to do what? 3 Enforcement and recovery What happens next, and is it proportionate to the risk? 4 Accountability Who owns agent policy, and do they have the authority to change it? 5 Standards and ecosystem Plugged into emerging identity and payment standards, or alone?
The Five Dimensions of Agent Trust Maturity. The levels are the vertical axis and the dimensions the horizontal; Figure 6 in the appendix is the two crossed.

The model is built to run alongside the frameworks security and risk teams already report against, not to replace them: visibility lines up with DETECT in the NIST Cybersecurity Framework, enforcement spans RESPOND and PROTECT, and accountability and standards integration sit with GOVERN in both CSF 2.0 and the NIST AI Risk Management Framework. A crosswalk appears in the appendix for readers who want to map this model onto a framework they already run.

CLASSIFICATION IS THE EXCEPTION, AND THAT IS THE POINT NIST's Cybersecurity Framework tells an organization to detect anomalies and respond to incidents. Its AI Risk Management Framework tells one to map context, measure risk, and manage it. Neither has a name for the function in between: deciding what a specific agent is actually trying to do, in the moment, before anything has gone wrong. Classification is the function the established frameworks have not yet named, and it is precisely where bot management tooling and agentic traffic keep missing each other.

What Are the Five Levels of Agentic Trust Maturity?

Level 1  Initial (Ad Hoc)

Agent traffic has no formal category inside the business yet, even if leadership has started asking questions about it. Operationally, it is still noise somewhere inside “bot traffic” or “suspicious traffic” more broadly, with no dedicated way to tell a legitimate AI checkout agent from a credential-stuffing bot, and often no one looking closely enough to draw that line either way.

Visibility: no dedicated telemetry for agent traffic; agent sessions are invisible inside aggregate reporting.

Classification: none. Traffic is evaluated only as human versus suspicious.

Enforcement: binary and reactive, applied inconsistently across endpoints, usually only after an incident. There is no remediation path for a wrongly blocked customer, and no way to unwind a wrongly allowed transaction. Anything caught goes to manual review.

Accountability: no named owner; it is common for fraud, product, and engineering to each assume someone else is watching.

Standards: not evaluated; the company is not tracking AP2, Trusted Agent Protocol, or similar identity standards.

Cost of staying at Level 1

Legitimate agent-driven revenue gets blocked as a false positive, and malicious agentic traffic gets through unnoticed until a chargeback spike or account-takeover wave forces a reaction.

Level 2  Repeatable (Reactive but Inconsistent)

Agent handling typically gets folded into whatever bot management tooling already exists, rather than built as its own function. The tooling was written to catch scripted bots, not agents mimicking human behavior on real devices and real IP ranges, so it misses the traffic it was never designed for. Response is repeatable within a single team, but not standardized company-wide.

Visibility: some agent-related telemetry exists, usually as a byproduct of legacy bot management, not purpose-built for agents.

Classification: basic allow or deny lists by user agent string or known crawler; typically no distinction between a self-disclosing partner integration and a non-disclosing consumer AI assistant.

Enforcement: the same binary block-or-allow logic as legacy bot rules; false positives tend to climb as legitimate agents get treated like adversaries. Recovery from a wrong call is manual and ticket-driven, handled case by case by support rather than by policy, and review capacity is already the bottleneck.

Accountability: owned by whichever team purchased the bot management tool, usually fraud or security engineering, with no cross-functional review.

Standards: aware of emerging agent identity standards, but not integrated with any of them.

Cost of staying at Level 2

Inconsistent treatment across endpoints, checkout blocks agents while search allows them, creates a confusing and sometimes revenue-losing experience. The company still cannot say with confidence how much of its traffic is agentic.

Level 3  Defined (Standardized Classification)

The company has a documented, company-wide definition of agent traffic populations and a policy for each, typically some version of self-disclosing good agents, non-disclosing good agents, and adversaries. The three are not evenly sized, and a policy that treats them as though they are will misallocate effort: disclosure has no ecosystem support behind it yet, so self-disclosing agents are a negligible share of live traffic today, while non-disclosing legitimate agents are where the volume sits and where it will stay for the foreseeable future. Fraud, product, legal, and engineering have agreed on what “authorized agent” actually means.

Visibility: a dedicated agent detection layer, behavioral signals, device intelligence, disclosure correlation, separate from generic bot rules.

Classification: a documented framework distinguishing self-disclosing good agents, non-disclosing good agents, and malicious adversaries, applied consistently across the business.

Enforcement: policy-driven, though still largely static; rules are set and revisited on a fixed cadence rather than continuously. A documented remediation path exists for false positives, though it is triggered by customer complaint rather than detected proactively, and review is still performed at human speed.

Accountability: a named cross-functional group, typically fraud plus product, owns agent policy and reviews it on a regular schedule.

Standards: actively piloting or adopting at least one agent identity standard, such as Web Bot Auth or AP2, for at least one channel.

Cost of staying at Level 3

The company can finally measure agentic traffic and treat legitimate agents as a channel rather than only a threat. The risk at this level is static, cadence-based policy lagging fast-moving agent behavior between review cycles.

Level 4  Managed (Quantitatively Controlled)

Agent trust is run like any other measured operational program. The company tracks agent traffic composition, intent distribution, and enforcement outcomes as ongoing metrics, and enforcement responds in real time rather than waiting for the next scheduled review.

Visibility: real-time dashboards showing agent-versus-human composition and intent distribution across critical flows such as login, checkout, and account creation, plus detection of the agent-to-human handoff, the point at which an agent that cannot complete a step returns control to the person it is acting for. That transition is observable, it is one of the few signals that survives the arrival of on-device agents, and most programs are not yet capturing it.

Classification: confidence-scored classification (for example, confirmed, likely, or suspected) rather than a single binary verdict, scored per action rather than per session, with continuous re-classification as behavior drifts inside a single visit.

Enforcement: a graduated spectrum, allow, monitor, challenge, throttle, hand over, block, mapped to intent tier and endpoint risk, not a single switch. Handing the decision back to the human is different in kind rather than degree, since it returns a legitimate customer to the flow instead of ending the session. Two different events share this vocabulary, and this report keeps them apart. Hand over is an enforcement action the platform takes, returning a decision to the human. The handoff is the agent's own move, when it gets stuck and gives control back to its user. The first is a control and belongs to enforcement. The second is a signal and belongs to visibility. Wrong calls in either direction are measured, with a defined path to reinstate a wrongly blocked customer and to unwind a wrongly allowed transaction. Review is partly automated, with humans handling exceptions rather than the queue.

Accountability: agent trust has a defined owner with a budget and measured outcomes, false positive rate, agent-attributed revenue, incident rate, reviewed at the executive level.

Standards: integrated with one or more payment-network agent protocols for verified transaction flows.

Cost of staying at Level 4

The company can quantify the cost of over-blocking and under-blocking, and can defend its agent policy with data. The limitation at this level is detection signals evolving more slowly than agent behavior, since improvement is still periodic rather than continuous.

Figure 5
LEVELS 1 TO 2: TWO OPTIONS Allow Block Inherited from legacy bot rules and applied inconsistently. The punishment fits the endpoint's original design, not the risk. LEVEL 3: SAME TWO OPTIONS, NOW POLICY-DRIVEN Allow Block Consistent across every endpoint, still reviewed on a fixed cadence rather than in real time. LEVEL 4 AND ABOVE: SIX OPTIONS, MAPPED TO INTENT AND ENDPOINT RISK Allow Monitor Challenge Throttle Hand over Block Verified good agent Confirmed adversary increasing friction, decreasing confidence
The Enforcement Spectrum. Handover returns the decision to the human rather than ending the session. Hand over is unfilled because it differs in kind rather than degree.

Level 5  Optimizing (Continuous, Adaptive Trust)

Agent trust is treated as a living system rather than a project. Detection improves continuously against new agent behavior, the company both contributes to and draws on shared threat intelligence, and agentic commerce is designed into the product experience rather than bolted onto the security stack after the fact.

Visibility: continuous, network-level visibility, often strengthened by shared detection signals across other companies facing the same agent traffic.

Classification: adaptive models raising or lowering confidence automatically as new agent frameworks appear, paired with deterrence matched to where the agent runs. Proof of work raises an attacker’s cost against a cloud-hosted agent, but does little against an agent running locally on a real device, which has real compute and will simply solve it. The effective deterrent for a local agent is a step it cannot complete without the human returning to the flow. That control works at the moment of the step rather than permanently, and this model treats the difference as a requirement rather than a footnote. Once a person has assisted an agent through a step up, the session carries authenticated state forward, so whatever the agent was authorized to reach stays reachable afterwards. A step up is therefore only as strong as the scope of the authority it confers, which is why Level 5 requires that authority to be bounded to the action in hand and re-established when the agent returns, rather than assumed to persist for the life of the session.

Enforcement: automated and continuously re-tuned; policy changes take effect in near real time with little to no change-management cycle. Recovery is designed in rather than bolted on: wrong calls feed back into classification automatically, and remediation reaches the affected customer without them having to ask. Detection and response run as automated loops matching the adversary’s iteration speed, with humans retained for ambiguity and accountability.

Accountability: agent trust strategy sits with product and revenue leadership as well as security, because agentic traffic is now a channel to grow, not only a risk to contain.

Standards: an active participant in standards development, FIDO Alliance working groups and network trust programs among them, helping shape the rules rather than only adopting them.

Cost of staying at Level 5

In principle, this is where agentic commerce becomes a competitive advantage rather than only a risk to contain, faster and more confident trust decisions should translate into a faster yes to legitimate agents.

Where Most Companies Actually Stand Today

Across the sourced data above, and consistent with what Arkose Labs sees across its own customer base, the market is concentrated at Level 1 (Initial) and Level 2 (Repeatable). Purpose-built agent trust tooling, and the payment-network standards it interoperates with, only began shipping in 2026, so the tooling required to reach Level 3 (Defined) has barely existed long enough for most companies to adopt it. Reused bot management tooling was never built to tell a legitimate agent from an adversarial one, which is why so many companies experience false positives against legitimate agents and false negatives against adversarial ones at the same time.

The pattern shows up by name in the field, not only in survey data. A review of companies across sectors found the same Level 1 accountability gap repeating itself: a European retailer running agentic commerce across eleven countries where security ownership sits with a different team entirely, an insurer with no policy for agentic payment automation, and a betting operator with no authorization framework at all (Gosschalk, BotBusters Summit, 2026). None of these are edge cases. They are what Level 1 looks like once agentic traffic arrives at a company that has never named an owner for agent policy.

Level 4 (Managed) and Level 5 (Optimizing) remain rare, concentrated among large payments, marketplace, and travel platforms with the scale to justify a dedicated agent trust program. For most companies, the realistic near-term goal is closing the gap between Level 2 and Level 3: replacing reused, tool-driven bot rules with a documented, cross-functional, three-population policy applied consistently across the business. This single move addresses the two failure patterns we see most often: blocking legitimate agent-driven revenue and missing adversarial traffic no longer shaped like a traditional bot.

Which Level Are You? Three Questions to Calibrate

Three questions to calibrate roughly where you sit. Each assesses a different dimension, and the point at which each one tips from one level to the next. Answer honestly rather than aspirationally:

Three calibrating questionsFull tool: 5 questions, then 10
Can you state, right now, what percentage of traffic on your checkout and login flows is agentic?
If the honest answer is “we don’t know,” you are at Level 1 on visibility.
VisibilityL1 vs L2
Can you distinguish agents helping your customers from agents harming them, using more than device or technical fingerprinting?
Behavioral, not just technical, evidence is what separates Level 2 from Level 3, and it is the question the consumer trust gap turns on.
ClassificationL2 vs L3
Does your enforcement give you more than two options, and when a decision turns out to be wrong, is there a defined path back for the customer?
Moving beyond binary block-or-allow is what separates Level 3 from Level 4.
EnforcementL3 vs L4
The full assessment scores each dimension separately, because most organizations sit at different levels across the five.
Take the assessment →

Moving Up the Curve

From Level 1 to Level 2: starts with measurement, not enforcement. Instrument agent detection separately from generic bot detection before writing a single new rule. You cannot set policy for traffic you cannot see.

From Level 2 to Level 3: is a policy exercise before it is a technology purchase. Bring fraud, product, legal, and engineering into one room to agree on a documented, three-population definition, self-disclosing good agents, non-disclosing good agents, and adversaries, and apply it consistently across endpoints, not just the ones already holding a security team’s attention.

From Level 3 to Level 4: requires trading static rules for a graduated, confidence-scored response. Replace the fixed review cadence with continuous re-classification, and replace binary block-or-allow with an enforcement spectrum, allow, monitor, challenge, throttle, hand over, block, so the punishment fits the risk rather than the endpoint’s original design. A challenge an agent can solve is a delay rather than a deterrent, which is why forcing the handover matters. Decide at the same time what the session is permitted to do after a human has assisted it, because the handover is what grants that authority and it is granted once. Measure wrong calls in both directions and build the path back before you need it.

From Level 4 to Level 5: is where the function changes ownership as much as it changes technology. Agent trust moves from a security cost center to a shared responsibility with product and revenue leadership, because by this stage the company is no longer only defending against agentic fraud, it is competing for agentic commerce revenue.

A Closing Perspective

This model is the lens Arkose Labs brings to every conversation about agentic traffic: detection alone answers only its first question, and a Level 1 or 2 company with excellent detection can still fail at classification, enforcement, or accountability. Arkose Agent Trust Manager was built around the three-population framework and the enforcement spectrum described in this report, allow, monitor, challenge, throttle, hand over, block, because a binary block-or-allow decision cannot serve a Level 3 or higher operating model. Companies today need to see every agent, understand what it is trying to do, and control the outcome accordingly.

This model explains the five failure modes Harvard Business Review identified more directly than first appears. An agent acting beyond what the customer authorized is a classification failure, and prompt injection is now its most common mechanism: an agent carrying instructions its customer never gave, inside a session that is otherwise legitimate, authorized and authenticated. That is precisely the case a binary automated-or-human verdict cannot see, and it is why classification has to run continuously through a session rather than once at its start. An agent mishandling sensitive information is an enforcement and standards failure. An agent failing with no path to recovery is an enforcement failure, and it is the one this report now tracks at every level. The remaining two, misunderstanding the product and misrepresenting the brand, are product and brand questions sitting outside this model, and we would not claim otherwise. But three of the five are squarely a maturity problem, which means each level is a claim about how many of them a company can prevent, and how quickly it notices when it does not.

The companies moving fastest from here will not be the ones blocking the most agentic traffic. They will be the ones able to tell, with confidence, which agents are respecting the rules of the road.

Find your level

The companion self-assessment places your organization on this scale in about 90 seconds: five questions, one per dimension. Ten more afterwards, if you want them, put three answers behind every level. It returns a level for each dimension and the specific next move from where you are today.

Take the assessment →

Frequently Asked Questions

What is the Agentic Trust Maturity Model?

A five-level framework, Initial, Repeatable, Defined, Managed and Optimizing, for how organizations detect, classify and control AI agents acting on customers’ behalf. It is assessed across five dimensions: visibility and detection, classification, enforcement, accountability and standards integration.

What are the five levels of agentic trust maturity?

Level 1, Initial, means agentic traffic is invisible inside general bot noise and nobody owns the problem. Level 2, Repeatable, means legacy bot rules get reused for agents inconsistently. Level 3, Defined, means a documented, cross-functional policy classifies agent populations consistently. Level 4, Managed, means agent trust runs as a measured program with real-time visibility and graduated enforcement. Level 5, Optimizing, means agent trust is a continuously improving capability rather than a security cost center.

How is agent classification different from bot detection?

Bot detection asks whether a session is human or automated. Classification asks a different question: once you know a session is an agent, is it helping your customer or working against them, and what is it actually trying to do? Legacy bot management answers only the first question, which is why non-disclosing legitimate agents and adversarial agents both get treated as bots today.

What is the difference between a hand over and a handoff in agent enforcement?

A handover is an enforcement action the platform takes, returning a decision to the human when an agent’s confidence or intent doesn’t clear the bar. A handoff is the agent’s own move, when it hits a step it can’t complete and returns control to the person it’s acting for. The first is a control. The second is a signal.

What industry standards apply to agentic commerce?

Google’s Agent Payments Protocol (AP2) uses signed Intent, Cart and Payment mandates to authorize an agent’s transaction scope. Mastercard’s Agent Pay and Verifiable Intent, and Visa’s Trusted Agent Protocol, add network-level verified agent identity on top of card rails. Google and Mastercard contributed AP2 and Verifiable Intent to the FIDO Alliance in May 2026, which is now developing them into an interoperable standard.

Appendix

A1. The maturity matrix in full

The matrix below is the entire model on one page, every dimension against every level, including what each level costs the business. It condenses in table form what the level sections describe in prose, and sits here so it can be referenced or lifted on its own.

Figure 6
LEVEL 1 Initial LEVEL 2 Repeatable LEVEL 3 Defined LEVEL 4 Managed LEVEL 5 Optimizing Visibility and detection None dedicated. Invisible in aggregate reporting. Byproduct of legacy bot management. Not purpose-built. Dedicated agent detection layer, separate from bot rules. Real-time dashboards by flow: login, checkout, signup. Continuous, network- level. Shared signals across companies. Classification None. Human versus suspicious only. Allow or deny lists by user agent string. No disclosure distinction. Three populations: disclosing good, non- disclosing good, adversary. Confidence-scored: confirmed, likely, suspected. Re-scored. Adaptive models plus deterrence matched to where the agent runs. Enforcement incl. recovery Binary, reactive, inconsistent. No remediation path. Same binary logic as bot rules. Recovery is manual, ticket-driven. Policy-driven but static. Remediation documented, complaint-triggered. Graduated spectrum by intent and risk. Wrong calls measured, path back. Automated, re-tuned continuously. Recovery designed in, proactive. Accountability No named owner. Each team assumes another is watching. Whichever team bought the tool. No cross- functional review. Named cross-functional group. Scheduled policy review. Defined owner with budget. Outcomes reviewed at exec level. Shared with product and revenue leadership, not security alone. Standards and ecosystem Not evaluated. Not tracking AP2 or Trusted Agent Protocol. Aware of emerging standards, integrated with none. Piloting at least one standard on at least one channel. Integrated with one or more payment-network agent protocols. Active in standards development. Shapes the rules. Cost of staying here Revenue blocked as false positives; fraud missed until a chargeback spike. Inconsistent treatment across endpoints. Still cannot say how much traffic is agentic. Static policy lags fast- moving agent behavior between review cycles. Detection signals may not evolve as fast as agents do. Improvement is periodic. Unproven. No company observed sustaining this long enough to confirm the payoff.
The Maturity Matrix: Every Dimension Against Every Level.

A2. Crosswalk to NIST CSF 2.0 and the AI Risk Management Framework

Most security and risk teams already report against a NIST framework. This crosswalk maps the dimensions in this report onto the NIST Cybersecurity Framework 2.0 and the NIST AI Risk Management Framework, so the model can be used alongside a framework you already run rather than instead of it.

This is a mapping of functions, not a claim of control-level compliance.

Table 2
THIS REPORT NIST CSF 2.0 NIST AI RMF FIT Visibility and detection DETECT MAP Direct. DETECT covers timely discovery, anomaly detection, and continuous monitoring. Classification No direct equivalent MEASURE, partially None. No established framework names the function of determining what a detected agent is trying to do. Enforcement RESPOND, with elements of PROTECT MANAGE Direct. RESPOND covers mitigation and containment; PROTECT covers access control. The enforcement spectrum spans both. Accountability GOVERN GOVERN Direct. Both frameworks place accountability, policy, and authority in GOVERN. Standards and ecosystem integration GOVERN, supply chain and external participation GOVERN Partial. NIST folds external standards participation into GOVERN. This report treats it separately because the standards layer is new and moving monthly. Recovery, within Enforcement RECOVER MANAGE, partially Partial by design. Recovery appears inside each level's Enforcement criteria rather than as a separate dimension.
Crosswalk to NIST CSF 2.0 and the NIST AI RMF.

On the classification gap

CSF’s IDENTIFY function concerns asset and risk inventory, not per-session judgment. The AI RMF’s MEASURE function concerns assessing system-level risk characteristics, not classifying live traffic in the moment. Neither maps to the question this report treats as central: what is this specific agent trying to do, right now, before anything has gone wrong. NIST’s Control Overlays for Securing AI Systems project includes single-agent and multi-agent use cases, but those overlays remain in development with publication projected for late 2026 into 2027. Separately, the National Cybersecurity Center of Excellence published a concept paper in February 2026 proposing to adapt existing identity and authorization frameworks for AI agents. The vocabulary for agentic control is still being written, which is why this report names the function rather than waiting for a standard to name it.

On terminology

This report uses Accountability rather than “governance” for the dimension mapping to NIST’s GOVERN function. The choice is deliberate. Accountability names what the dimension actually measures: who owns agent policy, and whether they hold the authority to change it. “Governance” has become a broad product category in this market, and buyers consistently report difficulty separating governance positioning from demonstrable capability. Readers running a CSF-aligned program should treat this dimension as their GOVERN function.

On recovery

NIST’s RECOVER function has no dimension of its own here. Recovery appears instead as a marker within Enforcement at each level, progressing from no remediation path at Level 1, through manual and ticket-driven at Level 2, documented but complaint-triggered at Level 3, measured with a defined path back at Level 4, to automatic and proactive at Level 5. Readers running a CSF-aligned program should expect to score recovery separately in their own reporting.

Sources

  1. commercetools, “Agentic Commerce Stats 2026: Enterprise Guide,” 2026. commercetools.com
  2. Visa, agentic commerce consumer forecast and shopping-task adoption data, 2026.
  3. MetaRouter, “Agentic Commerce Trends and Statistics for 2026,” 2026. metarouter.io
  4. The Wise Marketer, “Riskified Study Finds Consumers Aren’t Ready to Hand Over Control as AI Transforms Shopping,” 2026. thewisemarketer.com
  5. Harvard Business Review, Ali Furman, Ege Gürdeniz, Rima Safari, and Remzi Ural, agentic commerce and consumer trust analysis, February 2026.
  6. Gravitee, “State of AI Agent Security 2026 Report: When Adoption Outpaces Control,” 2026. gravitee.io
  7. Bank Director, “2026 Risk Survey: AI Exposes Threats, Knowledge Gaps,” sponsored by Baker Tilly, 30 March 2026. bankdirector.com
  8. Kevin Gosschalk, “Good agents, bad agents: from bot detection to agent trust,” BotBusters Summit, 18 August 2026.
  9. Eco, “AP2 Protocol Explained: Google’s Agentic Commerce Standard 2026,” 2026. eco.com
  10. FIDO Alliance, “Building the Trust Layer for Agentic Payments with AP2 and Verifiable Intent,” 2026. fidoalliance.org
  11. PYMNTS, “Google and Mastercard Contribute Agentic Commerce Standards to FIDO Alliance,” 2026. pymnts.com
  12. National Institute of Standards and Technology, The NIST Cybersecurity Framework (CSF) 2.0, NIST CSWP 29, February 2024. nvlpubs.nist.gov
  13. National Institute of Standards and Technology, AI Risk Management Framework (AI RMF 1.0) and Playbook. airc.nist.gov
  14. Cloud Security Alliance, research notes on the NIST AI Agent Standards initiative and the status of Control Overlays for Securing AI Systems, 2026. labs.cloudsecurityalliance.org