AI Governance & Compliance
The Capability-Authorisation Gap: Why Knowing What Your AI Agent Can Do Is Not Enough
Capability testing isn’t authorisation. The WEF identifies authorisation as AI governance’s real bottleneck. Here’s how to close the gap.
Published on


The Capability-Authorisation Gap: Why Knowing What Your AI Agent Can Do Is Not Enough
By Tahir Mahmood, Co-founder & CTO, OpenBox · Last updated 8 October 2026
In brief: The World Economic Forum’s May 2026 playbook with Capgemini argues that authorisation, not capability, is now the real bottleneck to large-scale AI agent adoption. Organisations are increasingly able to describe what their agents can do, but often struggle to say who approved what those agents are permitted to do, and under what conditions. Closing that gap means moving from capability documentation to a deployment-level authorisation practice built on the WEF’s own building blocks: enterprise agent guidelines and a deployment-level authorisation record, which can in turn underpin an auditable agent registry. |
The Billion-Dollar Question Everyone in AI Is Answering Wrong
Picture a mid-size insurer that has done everything by the book. Its claims-handling agent passed a full battery of red-team tests. Its benchmark scores are strong. Its model card is thorough. It is also running in production with no formal record of what it is allowed to do, who approved that scope, or when the approval expires. This is a hypothetical, but it illustrates exactly the authorisation gap the WEF identifies in early agent deployments.
The World Economic Forum, working with Capgemini, named why in its May 2026 playbook on AI agent adoption: organisations are increasingly able to describe what an agent is capable of, but often lack a consistent way to determine what it is authorised to do within a specific workflow. That gap between capability and authorisation is, in the report’s own words, the central challenge to large-scale adoption.
Enterprise AI programmes invest heavily in capability evaluation: benchmarks, red-team exercises, safety assessments, model cards. These are legitimate and necessary. They are also incomplete. They answer the question of what an agent can do. They do not answer what it is permitted to do in a specific workflow. The WEF’s playbook, the third in a three-part series on AI agents, states the distinction without hedging: AI agents are advancing faster than governance, and authorisation, not capability, is now the critical bottleneck.
An agent with carefully evaluated capabilities but no authorisation record is not governed. It is described. The failure sits in the permission space, not the capability space: authorisation is fundamentally an organisational decision, but one that has to be technically enforced, not just written down.
Defining the Gap: Capability vs. Authorisation
Capability and authorisation answer different questions, and treating them as one is where AI agent authorisation programmes break down. Capability is architectural: the range of actions an agent’s model, tools, and data access make technically possible, established through evaluation and testing. Authorisation is organisational: the explicit, context-specific set of actions an agent has been approved to take in a defined workflow, by named accountable owners, under defined constraints.
Dimension | Capability | Authorisation |
|---|---|---|
What it answers | “What can this agent do?” | “What is this agent permitted to do, here?” |
How it is established | Evaluation, benchmarking, red-teaming | Named owner sign-off against enterprise guidelines |
What it produces | Model card, evaluation report, benchmark score | A deployment-level authorisation record (an ACAP) |
Who owns it | Developers and evaluation teams | Deployment owner, with risk, compliance, and SME sign-off |
Changes with context? | Model-level evaluated capability travels with the model; deployment context determines how that capability is exposed, constrained, and authorised. | Yes. The same agent can carry a different risk profile in a different deployment context. |
A fully evaluated agent with no authorisation record is capable, not governed. The permissions that make an agent valuable in one workflow, such as external communication, financial data access, or write permissions, make it dangerous in another without an explicit authorisation boundary attached to that specific deployment. The WEF’s report makes this concrete: the same agent can carry very different risk and governance implications depending on which organisational or technical boundaries it crosses.
This becomes sharper in multi-agent systems. The playbook specifies that authorisation is non-additive: a downstream agent operates only under the intersection of its own permissions and those of the agent that invoked it, and an orchestrating agent cannot delegate authority it does not itself hold. We’ve written separately about how authorisation propagates, or fails to, across a multi-agent pipeline.
The Employee Analogy: Why We Already Know How to Do This
The WEF compares onboarding an agent to onboarding a new employee: the process, in its words, involves assessing capabilities, assigning responsibilities, granting access, defining supervision, and adjusting the scope of action as performance and trust evolve. The governance gap this article describes appears when capability assessment is completed without equally explicit responsibility assignment, access control, supervision, and scope management, the other four steps in that same list.
In human onboarding, a capable candidate does not start working unsupervised on day one. A manager defines the role. A named supervisor signs off on access. The scope of responsibility expands only as trust is earned, and each step has a named owner and a documented decision behind it.
When that gap is left open, the leap from evaluation to production has nothing in between: the agent is benchmarked, deployed, and then, in effect, hoped about. The explicit responsibility assignment, the defined approval authority, and the structured trust expansion that make human onboarding safe have no equivalent step in that path to production.
The WEF is explicit that the comparison is structural only: agents lack clear legal status, moral responsibility, and reputational incentive to stay within their mandate, and governance has to be built to offset those gaps. That disanalogy cuts one way. It makes authorisation more necessary for agents than for employees, not less. A human who exceeds their mandate faces consequences that create a natural brake. An agent does not, until a human has built one in.
What the Capability-Authorisation Gap Costs Organisations
Leaving the gap open carries four costs: audit exposure, accountability fragmentation, shadow agents, and an inability to scale with confidence. Each compounds the others.
Audit and regulatory exposure
When an auditor or regulator asks who approved what an agent is doing, “the model was evaluated as capable” answers a different question than the one being asked. The EU AI Act's Article 14 requires that high-risk AI systems be designed so they can be, in the regulation's own wording, “effectively overseen by natural persons during the period in which they are in use.” That obligation will apply to standalone high-risk systems under Annex III from 2 December 2027, after the EU's Digital Omnibus deferred the original 2 August 2026 date. The amendment changes the application date; it does not remove the Article 14 requirement itself, and capability documentation alone will not show who is exercising that oversight, or on what authority, once the date arrives.
Accountability fragmentation
Agent guidelines exist, in the WEF's own framing, so that deployment decisions are comparable, so that responsibility does not fragment across teams, and so that agent authority expands based on evidence rather than ad hoc practice. Without them, when something goes wrong, it becomes harder to identify a single accountable function and the evidence behind the agent's approved scope. We've looked at the accountability gap when no verifiable authorisation record exists in more detail elsewhere.
Shadow agent risk
Low-code and agent-building tools make creating an agent easy. Authorizing one formally is harder, and the two can move at different speeds. Local experiments can grow in scope without ever entering a governance process. Call these shadow agents: capable, deployed, and never authorised. The WEF’s own trigger for when this has to stop is specific: an agent should enter formal governance when it begins to “interact with enterprise systems, operate across teams or perform consequential events.” This is close to the shadow IT pattern this creates.
Scale failure
Without a consistent authorisation practice, every new deployment becomes its own governance negotiation. Authority expands by ad hoc practice rather than by evidence-based promotion gates, which means the organisation cannot scale its agent estate with confidence that governance is keeping pace. The alternative is what the next section describes: a registry built from standardised authorisation records, which is how the WEF frames the shift from isolated pilots to a governed portfolio of agents.
Closing the Gap: From Capability Documentation to Deployment Authorisation
The WEF’s playbook itself names three building blocks: agent guidelines, the ACAP, and the adoption life cycle that builds and maintains it. For practical implementation, that maps to three linked artefacts built in order: enterprise agent guidelines, a deployment-specific authorisation record, and a registry that rolls those records up into one auditable inventory, the last of which the WEF describes the ACAP as able to underpin. Skipping a step leaves the next one without a foundation.
Step 1: Agent guidelines
A shared, enterprise-level policy that defines terminology (autonomy, authority, consequential events, boundaries), decision rights, and deployment criteria. This is set once, by leadership, and applied to every deployment that follows. It is the governance foundation everything else builds on.
Step 2: The ACAP
For each specific deployment, the WEF’s playbook calls for a context-specific authorisation record, the Agent Capability and Authorisation Profile, defining what the agent is permitted to do, who approved it, what evidence justifies that permission, and when it must be reviewed again. This is the artifact that actually closes the gap: not a benchmark score, but a signed, reviewable record tied to one workflow.
Step 3: The agent registry
A registry is what the ACAP enables at scale: a centralised, standardised view of every deployed agent, its authorised scope, and its accountable owner. It makes the governed agent estate visible, and it is what lets an auditor find evidence instead of asking around for it. For teams building this out, a week-by-week plan for finding and governing agents nobody approved is a useful place to start.
None of this is quick. Agent guidelines call for cross-functional agreement on terminology and decision rights before a single deployment is authorised. The WEF’s model calls for an ACAP to carry named sign-off from deployment owners and risk and compliance functions, with the relevant subject matter experts involved for consequential events. Building this is organisational work, not a configuration setting.
An authorisation record is also not the end of governance. It is the starting point. The WEF’s own model treats monitoring, incident response, and periodic re-authorisation as ongoing phases, not one-time paperwork. Closing the gap once does not keep it closed.
Gap Diagnostic: Six Questions Your Team Should Be Able to Answer
Use this as a short self-assessment. If the answer to any of the following is “no” or “I’m not sure,” the capability-authorisation gap is present in your organisation’s deployment programme.
1. For each AI agent currently in production, can you name the specific person who approved what it is permitted to do in its workflow?
2. Is there a written record specifying which agent actions require human approval before execution?
3. Does your organisation have a defined criterion for when a bottom-up agent experiment must enter formal governance?
4. Can you produce a list of every AI agent currently deployed, its authorised scope, and its accountable owner?
5. If an agent’s capabilities expand through a model update or a new tool, is there a defined re-authorisation process?
6. Does your current governance documentation distinguish between what an agent can do and what it has been authorised to do in context?
How OpenBox Approaches This Gap
OpenBox's Trust Lifecycle separates these two questions into two explicit phases. Assess produces a quantitative Risk Profile Score for each agent, weighted across Base Security, AI-Specific, and Impact factors, which feeds into an overall Trust Score. Authorize is where guardrails, policies, and Behavioral Rules are configured to govern what an agent can actually do at runtime, enforced as one of five decisions: ALLOW, CONSTRAIN, REQUIRE_APPROVAL, BLOCK, or HALT.
This does not replace the organisational sign-off the WEF describes. A named deployment owner still has to decide what an agent is permitted to do in a given workflow, and risk or compliance functions still have to approve that decision. What Assess and Authorize provide is somewhere for that decision to live: a risk score and a set of enforced, auditable runtime controls the authorisation decision can attach to, rather than a policy document that nothing downstream actually checks.
Capability evaluation is not going away, nor should it. But it answers a question that authorisation does not, and conflating the two is how a well-tested agent ends up running in production with nobody able to say who let it do what it is doing. The organisations that start treating authorisation as seriously as they already treat capability will be the ones who can put a name next to every agent’s permission, not just a benchmark score beside its name.
Frequently Asked Questions
What is a consequential event in AI agent governance?
A consequential event is an agent action with significant legal, financial, security, safety, ethical, reputational, or customer-facing impact, inside or outside the organisation. The WEF’s playbook treats these as the specific checkpoints an authorisation record must assign an approver and an escalation path to, rather than giving every action the same level of oversight.
What does “authorisation is non-additive” mean in multi-agent systems?
In a multi-agent pipeline, a downstream agent operates only under the intersection of its own authorised permissions and those of the agent that invoked it. The WEF’s playbook is explicit that an orchestrating agent cannot delegate authority it does not itself hold, so permissions narrow as a task passes between agents; they never expand.
Which roles should be involved in authorizing an AI agent deployment?
The WEF’s playbook names deployment owners, developers and engineering teams, subject matter experts, risk, compliance and legal functions, human supervisors, and HR or change-management functions, each with a specific, named decision right across the agent’s lifecycle. No single function owns authorisation alone.
Sources World Economic Forum & Capgemini, “AI Agents in Action: A Playbook for Trusted Adoption, Authorization and Scaling” (May 2026). reports.weforum.org. Accessed 8 October 2026. OpenBox, “Assess (Phase 1).” docs.openbox.ai/trust-lifecycle/assess. Accessed 8 October 2026. OpenBox, “Authorize (Phase 2).” docs.openbox.ai/trust-lifecycle/authorize. Accessed 8 October 2026. European Commission, “Enforcement of the AI Act” (Annex III date) and Article 14 (human oversight). digital-strategy.ec.europa.eu/.../enforcement-ai-act. Accessed 8 October 2026. EU AI Act Digital Omnibus, Regulation (EU) 2026/1744 (in force 27 July 2026), official text. eur-lex.europa.eu/eli/reg/2026/1744/oj/eng. Accessed 8 October 2026. |

