AI Governance & Compliance
The Agentic Organisation: From AI to Org Redesign
Most companies are running AI pilots. Few redesign how work gets done. Learn what changes when AI stops recommending and starts acting, and what governance must follow.
Published on


The Agentic Organisation: From AI Experimentation to Organisational Redesign
Most companies are running AI experiments. A small number are doing something categorically different. Understanding the gap is the beginning of capturing the value.
By Tahir Mahmood, Co-founder & CTO, OpenBox · Last updated 2 October 2026
AI is everywhere. The results are not
Consider the paradox McKinsey’s senior partner Alexis Krivkovich described in April 2026: companies expect massive transformations from AI and have invested with that mindset, yet more than 80 percent of companies say they are not yet seeing impact on the bottom line from those investments (McKinsey, “AI is everywhere. The agentic organization isn’t yet,” April 2026). The number is striking because it is not measuring scepticism about AI; it is measuring the experience of organisations that have already committed to it.
The logical question that follows is: why? If the technology genuinely works, why does widespread experimentation not produce widespread results? McKinsey’s September 2025 research on the agentic organisation offers a direct answer. The organisations McKinsey identifies as capturing scalable value are those that have redesigned how work gets done around AI, rather than those that have simply layered the most AI tools onto existing processes. Those are different things, and the difference matters more than it might initially appear.
The distinction that matters: adding AI versus redesigning around it
The most common enterprise AI deployment pattern is additive: an AI assistant reduces the time it takes to draft a document; a model screens candidate CVs before a human reviews them; a chatbot handles first-contact customer queries. Each of these is a genuine improvement. None of them is an organisational redesign.
McKinsey’s research draws a sharper line. According to the September 2025 article, the best use cases for AI, those enabling scalable impact, are where a workflow is reimagined in its entirety, because those workflows typically cut across multiple teams and areas. Rather than point solutions, where AI accomplishes one task better and faster, organisations are starting to ask: ‘How do I take something like insurance underwriting and rethink that end to end?’
The distinction is not about the sophistication of the underlying model. It is about whether AI is being inserted into a system designed for humans, or whether the system itself has been rebuilt to accommodate AI as a first-class participant. In the first case, AI might save hours. In the second case, it potentially changes what is possible in the first place.
McKinsey frames this as two different relationships between humans and AI agents. In one model, humans perform steps and AI assists with individual tasks within those steps. In the other, agents execute the bulk of the process while humans operate above the loop: steering and validating outcomes rather than performing the work itself. The operational implications of the second model are substantially different from the first, including in what governance is required.
What ‘agentic’ actually means (and what it does not)
‘Agentic’ is used loosely in enterprise technology conversations, and that looseness matters for anyone making governance or investment decisions. In McKinsey’s framing, an AI agent is not simply a chatbot or a model that produces outputs. An agent is a system that can plan, make decisions across steps, use tools, and execute workflows towards a goal. As McKinsey partner Rich Isenberg stated in March 2026:
In McKinsey’s March 2026 episode on trust and agentic AI, Rich Isenberg argued that giving AI systems agency is not a product feature; it is a transfer of decision rights. The governing question for organisations, he said, shifts from whether the model is accurate to who is accountable when it acts. (McKinsey, “Trust in the age of agents,” March 2026)
McKinsey’s September 2025 research describes AI agents operating along a spectrum of increasing complexity: from tools that augment existing activities, to end-to-end workflow automation, to AI-first systems where agents are the operating default. In the most advanced deployments, McKinsey notes that physical AI agents are also emerging, covering drones, self-driving vehicles, and early humanoid robots, though these remain at an earlier stage of enterprise deployment than virtual agents.
An agentic organisation, in McKinsey’s definition, is one built around five pillars: business model, operating model, governance, workforce, people, and culture, and technology and data. It is not simply an organisation that uses AI agents. It is one that has structured itself so that networks of human and AI agents collaborate to deliver outcomes, with humans positioned above the loop for strategic oversight. Only
1 percent of organisations acted as a decentralised network of the kind McKinsey associates with this emerging paradigm, according to its September 2025 research. 89 percent remain in industrial-age operating models, and 9 percent have reached the agile or product-platform models of the digital era (McKinsey, September 2025). The gap is large.
Why most organisations are not there yet
The barriers most organisations face are not primarily technological. McKinsey identifies several structural challenges in the transition to agentic operating models.
Workflows were not designed for agents. Traditional processes assume handoffs between people. They are built around the cognitive limits, working hours, and error patterns of human workers. An AI agent does not share those limits; it also does not share the contextual judgment humans bring to edges, exceptions, and situations not covered by the written process. Redesigning a workflow for agent participation requires mapping those judgment points explicitly, not adapting the existing process around a new tool.
Roles need fundamental reshaping. McKinsey’s research found that 75 percent of roles need fundamental reshaping right now (Krivkovich, April 2026). Not replacement: reshaping. The question is not whether AI replaces a role but whether the tasks within that role change enough that the person needs a different skill set, different measures of success, and different management support. McKinsey’s research found this type of role-level analysis remains incomplete in most organisations.
Governance has not kept up. A common enterprise governance pattern was designed around models that produce outputs for humans to review. When an agent executes a workflow, the governance surface changes. McKinsey’s research notes that the scale of agentic adoption will ultimately be capped by how much oversight capacity humans can provide, making governance itself a potential bottleneck to productivity. Organisations that invest in capability while underinvesting in governance are not scaling agentic AI; they are accumulating unpriced risk.
Talent, culture, and incentives lag. Change management in the agentic era is, as Krivkovich observed, “no longer an episodic thing. It is a perpetual state.” Organisations accustomed to periodic transformation programmes are being asked to build a capacity for continuous change, which is a different organisational competency.
The pilot-to-production gap is real. McKinsey’s Rich Isenberg described the common failure mode: most organisations can get one or two use cases running by pulling in their best people and having a debate. What does not scale is the governance and infrastructure around those pilots. The organisations that have moved fastest, Isenberg observed, are those that have invested in making governance a repeatable product rather than a bespoke committee process.
When AI moves from recommending to acting, the governance problem changes
This is where the enterprise AI conversation becomes operationally serious. A model that drafts a document and waits for a human to review and send it creates one category of risk. A model that drafts, classifies, routes, and sends the document based on its own assessment of what is appropriate creates a different category.
McKinsey reported that 80 percent of organisations surveyed had encountered risky behaviour from AI agents, citing a May 2025 SailPoint Technologies global survey of security, IT professionals, and executives. Among the failure modes Isenberg discussed: agents propagating a flawed decision downstream faster than human oversight could catch it, and agents behaving differently when context changed in ways the original policy did not anticipate. In his framing: “Agent risk isn’t just about wrong answers; it’s wrong answers at scale. What executives need to keep in mind is that the scariest failures are the ones you can’t reconstruct, because you didn’t log the workflow.”
A common AI governance pattern concentrates much of its effort before deployment: documentation, risk classification, evaluation, and approval. Established frameworks such as the NIST AI RMF already include post-deployment monitoring and response as lifecycle components, but in practice many enterprise governance programmes weight the pre-deployment activities most heavily. That model remains relevant. But an agentic system changes the operational surface after deployment in ways a pre-deployment review cannot fully anticipate. Consider the difference between these two scenarios:
Traditional AI model | Agentic model |
|---|---|
Input arrives | Goal is set |
Model produces output | Agent plans and selects steps |
Human reviews output | Agent calls tools and takes actions |
Human decides whether to act | Monitoring observes behaviour |
Action is performed | Feedback loop adjusts further action |
In the governance model described above, the human review step serves as the primary governance moment for traditional AI use. In an agentic governance architecture, enforcement is distributed across every action the agent takes rather than concentrated at a single review point. Pre-deployment risk classification determines what the agent is permitted to do; runtime governance determines whether it is doing exactly that, and stops it when it is not.
This is explored further in “Governance Starts Where Monitoring Ends” and “Runtime Enforcement Comes Before Compliance”, which document why observability after the fact is not a substitute for enforcement before an action executes.
The accountability question nobody has answered cleanly
Accountability in human organisations is built around a traceable chain: a person authorised an action, performed it, and can be held responsible for the outcome. Agentic systems strain that chain in at least four ways.
Separated authorisation and execution. A human sets a goal and gives the agent authority to act. The agent selects the execution path. If the outcome causes harm, the question of who is accountable depends on whether the harm was within the scope of what was authorised, whether the agent behaved within its defined boundaries, and whether those boundaries were adequately specified. None of those questions has a ready or automatic answer when governance was not built into the deployment from the start, and existing contractual, regulatory, or organisational frameworks may not map cleanly onto actions that no individual person performed.
Changing context after authorisation. A human approves a system at deployment time. The system later acts in a context that has changed: different data, different permissions, different external conditions. The original approval may no longer accurately represent what the system is doing. As Isenberg noted, without upgrading inventory, identity management, or observability, organisations that assume they have risk covered are “not scaling agents; they’re scaling unknown risk.”
Multi-agent chains. When an agent delegates to another agent, and that agent in turn calls third-party tools or APIs, the chain of accountability stretches across systems that may have different owners, different policies, and different logging practices. McKinsey describes, as an illustrative scenario, what would happen if a single data poisoning attack affected agents trained on the same internal knowledge: failures could propagate across operations, finance, and customer interactions simultaneously.
The absence of a verifiable record. If an action cannot be reconstructed end to end, demonstrating accountability becomes materially harder. This matters legally, operationally, and in terms of organisational trust. The “AI Oversight Boards Need Better Evidence” piece documents why logs alone are insufficient: what effective governance needs is evidence that is retained, attributable, and fit to support a finding, not merely a record of what the system emitted.
The Anthropic examples cited by McKinsey were simulated scenarios, not documented production incidents. However, the governance implications they illustrate are operational rather than hypothetical: they describe failure modes that become possible when agents can act without sufficient enforcement boundaries, and they have structural parallels in real-world agentic deployments. As McKinsey’s research on agentic governance states: “Ultimately, the scale of agentic adoption will be capped by how much oversight capacity humans can provide, making governance itself a potential bottleneck to productivity.”
Runtime governance versus pre-deployment governance
Pre-deployment governance is necessary. Risk classification, documentation, evaluation, and approval processes establish what a system is permitted to do and who has authorised it. Some of these activities are required for specific systems under applicable regulatory frameworks; others represent established risk-management practice regardless of regulatory status.
Agentic systems create a stronger need for controls that remain active during execution. The NIST AI Risk Management Framework (NIST AI RMF 1.0, released January 2023), which is designed for voluntary adoption and is currently being revised as part of the White House AI Action Plan, identifies monitoring and evaluation as components of responsible AI deployment throughout the AI lifecycle, not only at the pre-deployment stage. The Generative AI Profile (NIST AI 600-1, released July 2024) extends this analysis to the specific characteristics of generative AI.
The EU AI Act, which entered into force in August 2024 and became applicable on 2 August 2026, with some provisions subject to different timelines (with general-purpose AI model obligations applicable from August 2025), places distinct obligations on providers and deployers of high-risk AI systems: providers must maintain post-market monitoring systems, while deployers are required to ensure human oversight and to monitor the operation of the system and act on identified risks. For Annex III high-risk use cases, these obligations apply from 2 December 2027, following the timeline extended under the AI Omnibus regulation that entered into force on 27 July 2026. The Act does not define “agentic AI” as a risk category, but it does place obligations on systems that operate with degrees of autonomy that are relevant to agentic deployments. Organisations should read the legislation directly and seek qualified legal advice on which obligations apply to their specific systems (European Commission, “AI Act,” last updated 3 August 2026).
The distinction between pre-deployment and runtime governance is explored in depth in “Logging Is Not AI Governance” and “AI Agent Governance Cannot Be Optional”. The core point is that an agent can behave differently depending on the tools available to it, the external data it encounters, the model or provider in use, and the permissions it has been granted. Governance that exists only before deployment cannot adapt to those variations at runtime.
What an agentic workflow might look like in practice
McKinsey’s research describes a case from the American Arbitration Association (AAA) in which agents were trained on closed case files to review dispute submissions, assemble timelines, analyse both sides of a case, and produce a summary decision, with a human arbitrator retaining final judgement. This is a useful example because it illustrates the human-above-the-loop model rather than human-in-the-loop: the agent completes substantial process steps, and a human reviews and approves the outcome.
The following comparison illustrates the difference between a traditional and an agentic approach, presented as a conceptual model rather than a prescriptive recommendation. It uses claims management as the example, but the structure applies across professional services, finance, compliance operations, and similar knowledge-work domains.
Stage | Traditional claims workflow | Agentic claims workflow |
|---|---|---|
Intake | Employee receives and logs the claim | Agent receives, classifies, and validates the claim against policy |
Investigation | Employee researches precedents and gathers evidence | Agent retrieves authorised documents, runs initial fact analysis |
Recommendation | Employee drafts a recommendation and submits for review | Agent produces a structured recommendation with supporting evidence |
Approval | Manager reviews and approves or returns for revision | Human reviewer approves at a defined checkpoint; agent awaits confirmation |
Resolution | Employee executes the decision and records outcome | Agent executes permitted actions and records a cryptographically attested audit trail |
Oversight | Periodic audit by compliance team | Continuous runtime monitoring; alerts on anomalous behaviour |
Two aspects of the agentic column deserve attention. First, in the governance design McKinsey describes, the human review step does not disappear; it moves. Rather than reviewing every intermediate step, a human reviews at the checkpoint where the consequences of proceeding are material. This holds for workflows designed with that checkpoint in place; not every agentic deployment will follow the same pattern. Second, in the governance model being described, the audit trail is not a log file to be reviewed later. It is a continuously produced record of every decision and action, available in real time and retained in tamper-evident form. This represents a governance design choice, not a universal standard; the appropriate form of evidence will depend on the system, its risk classification, and applicable obligations. The practical difference is between an organisation that can explain what happened and one that must reconstruct it from partial records after something goes wrong.
Building the governance infrastructure before agents gain authority
McKinsey’s Rich Isenberg offered a practical framework for boards and senior leaders: five questions that any organisation should be able to answer with precision before claiming its agentic AI is under control. Paraphrasing from the March 2026 episode:
1. Do we have a complete inventory of agents and their owners? Agents that are not inventoried cannot be governed. An AI agent governance guide published by OpenBox notes that this is the starting point: not auditing agent behaviour, but knowing what agents exist and who is accountable for each one.
2. How is autonomy tiered by risk? Not all agents carry the same risk. An agent that drafts internal summaries operates in a different risk category from one that approves financial transactions or accesses production infrastructure. A risk-tiered approach to autonomous action is sound risk management and, in some jurisdictions, aligns with the direction regulators are taking: the EU AI Act classifies AI systems by risk level and use case, and this risk-based logic supports tiered governance of autonomous action as a reasonable design principle, even where no specific provision mandates the exact form.
3. Do agents operate under verified identities with least-privilege access? If an agent can access more than it needs for its defined task, it can cause more damage than was intended when it behaves unexpectedly. The OpenBox Trust Lifecycle addresses this through the Assess and Authorize stages, establishing an agent’s Risk Profile before it is permitted to operate and configuring the policies and guardrails that constrain its actions.
4. Can decisions be reconstructed end to end? An audit trail that cannot be used to reconstruct what happened is not a governance asset; it is an archive. This requires that the right events are logged with the right context, retained in a form that is attributable, and accessible without requiring forensic work. OpenBox Compliance and Audit documentation describes the governance event structure that makes this possible, including cryptographic attestation per session.
5. Is there a real rollback plan if something goes wrong? Governance is not only about preventing bad outcomes; it is about containing them when prevention fails. This requires that someone has mapped what actions are reversible, which are not, and what happens to pending operations if a session is terminated.
These five questions surface a distinction between what McKinsey calls “centralized governance with federated execution”: a central function that sets policies, maintains oversight, and defines risk thresholds, while individual business areas operate their own agents within those boundaries. In McKinsey’s proposed model, neither pure centralisation nor pure federation delivers adequate control: the recommended approach centres governance while distributing execution.
What runtime governance looks like when it is built in from the start
OpenBox, an AI agent governance platform designed for enterprises deploying agents in production, provides one approach to the infrastructure problem this article describes. Rather than treating governance as a review process, OpenBox wraps existing agents with a Trust Lifecycle: Assess, Authorize, Monitor, Verify, Adapt (docs.openbox.ai).
At deployment, the platform assesses the agent’s Risk Profile and produces a Trust Score, a 0-100 metric calculated as (Risk Profile Score x 40%) + (Behavioral Score x 35%) + (Alignment Score x 25%). The Trust Score determines which Trust Tier the agent operates in, which in turn determines how much autonomy it is granted and what level of human review is applied to its actions.
At runtime, every agent action is evaluated against one of five governance decisions: ALLOW, CONSTRAIN, REQUIRE_APPROVAL, BLOCK, or HALT. CONSTRAIN permits an action to proceed with active restrictions applied, such as rate limits or content guardrails, making it a graduated response positioned between outright permission and a full pause for human review. These follow a fixed precedence: HALT overrides everything, followed by BLOCK, then REQUIRE_APPROVAL, then CONSTRAIN, then ALLOW. When a REQUIRE_APPROVAL verdict is returned, the action pauses and a human reviewer in the approvals queue decides whether to proceed. When a HALT verdict fires, the entire agent session is terminated immediately. These are not soft suggestions; they are enforced at the middleware layer, outside the agent’s own discretion (docs.openbox.ai/core-concepts/governance-decisions).
The Verify stage produces cryptographically attested audit evidence per session, meaning that each session’s governance events are signed in a form that makes post-hoc alteration detectable. In the Adapt stage, Trust Scores are updated continuously based on observed behaviour, so an agent that has been compliant over time can earn greater autonomy, while one that has triggered violations faces tighter controls automatically.
This is the kind of infrastructure McKinsey’s governance research describes as necessary for organisations that want to scale agentic adoption rather than merely pilot it. The question the “Meaningful Human Control Must Be Enforceable” article poses is precise: human oversight is not meaningful if it is advisory. It becomes meaningful when it is built into the pathway an agent must travel before it acts.
Frequently asked questions
What is an agentic organisation?
An agentic organisation, in McKinsey’s definition from its September 2025 research, is one in which humans and AI agents work side by side at scale, with humans operating above the loop for strategic oversight and agents executing workflows end to end. It is built around five pillars: business model, operating model, governance, workforce, people, and culture, and technology and data. Being agentic is not simply a matter of using AI agents; it involves redesigning how work is structured around those agents.
What is the difference between AI adoption and agentic organisational redesign?
AI adoption adds AI to an existing organisation: a tool improves a task, a model assists a process, an assistant reduces a workload. Agentic redesign changes the underlying process itself: workflows are rebuilt around human-agent collaboration, roles are redefined, and governance structures are updated to match. McKinsey’s research distinguishes point solutions, where AI performs one task better, from end-to-end workflow reimagination, where the entire process is redesigned with AI as a participant.
What happens to governance when an AI agent takes actions rather than producing outputs?
The governance surface expands. A model that produces an output for human review requires governance at the review stage. An agent that plans, selects tools, and executes actions creates a need for governance across the execution path: what the agent is authorised to access, what actions are permitted, when a human must approve, what happens when a boundary is crossed, and how every decision is recorded. Pre-deployment risk classification is necessary but not sufficient; runtime controls and verifiable audit evidence can provide additional assurance over actions taken during execution.
Does the EU AI Act apply to AI agents?
The EU AI Act applies to AI systems based on their use case and risk level under the Act’s risk classification framework, rather than on whether they are described as “agentic.” Some agentic deployments will fall within high-risk categories under Annex III; others will not. The Act does not define “agentic AI” as a distinct risk category. High-risk obligations under Annex III apply from 2 December 2027, following the timeline extension in the AI Omnibus regulation that entered into force on 27 July 2026. GPAI model obligations have applied since August 2025. Organisations should assess their specific deployments against the Act’s risk classification framework and seek qualified legal advice.
Where does OpenBox fit into this?
OpenBox, an AI agent governance platform (openbox.ai), provides runtime governance infrastructure for organisations deploying agents in production. It wraps existing agents, including those built on LangGraph, LangChain, Mastra, Temporal, CrewAI, and similar frameworks, with a Trust Lifecycle covering Assess, Authorize, Monitor, Verify, and Adapt. It enforces five governance decisions (ALLOW, CONSTRAIN, REQUIRE_APPROVAL, BLOCK, HALT) at runtime, maintains a Trust Score per agent, and produces cryptographically attested audit evidence per session. It is designed to address the governance gap that emerges when agents move from pilots to production.
Sources |
|---|
McKinsey & Company, “AI is everywhere. The agentic organization isn’t yet,” Podcast, April 2, 2026. https://www.mckinsey.com/capabilities/people-and-organization/our-insights/ai-is-everywhere-the-agentic-organization-isnt-yet, accessed 15 September 2026. |
McKinsey & Company, “The agentic organization: Contours of the next paradigm for the AI era,” Article, September 26, 2025. https://www.mckinsey.com/capabilities/people-and-organization/our-insights/the-agentic-organization-contours-of-the-next-paradigm-for-the-ai-era, accessed 15 September 2026. |
McKinsey & Company, “Trust in the age of agents,” Podcast, March 5, 2026. https://www.mckinsey.com/capabilities/risk-and-resilience/our-insights/trust-in-the-age-of-agents, accessed 15 September 2026. |
National Institute of Standards and Technology, “AI Risk Management Framework,” last modified August 13, 2026. https://www.nist.gov/itl/ai-risk-management-framework, accessed 15 September 2026. |
European Commission, “AI Act,” last updated 3 August 2026. https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai, accessed 15 September 2026. |
OpenBox, “Trust Lifecycle,” docs.openbox.ai. https://docs.openbox.ai/trust-lifecycle.md, accessed 15 September 2026. |
OpenBox, “Governance Decisions,” docs.openbox.ai. https://docs.openbox.ai/core-concepts/governance-decisions.md, accessed 15 September 2026. |
OpenBox, “Trust Scores,” docs.openbox.ai. https://docs.openbox.ai/core-concepts/trust-scores.md, accessed 15 September 2026. |
OpenBox, “Compliance & Audit,” docs.openbox.ai. https://docs.openbox.ai/administration/compliance-and-audit.md, accessed 15 September 2026. |

