AI Governance & Compliance
How to Assess and Tier Agentic AI Use Cases by Risk Before You Deploy
Learn to apply a two-axis risk framework to AI agent deployments, covering impact, reversibility, and autonomy, with enterprise examples from Dayos and OCBC.
Published on


How to Assess and Tier Agentic AI Use Cases by Risk Before You Deploy
An operational approach to pre-deployment risk assessment, drawing on the IMDA MGF for Agentic AI’s likelihood-and-impact framework, with enterprise examples from Dayos, OCBC, and MSD (v1.5, May 2026).
By Tahir Mahmood, Co-founder & CTO, OpenBox · Last updated 21 September 2026
Why Risk Tiering Is the Starting Point for Agentic AI Governance
The Deployment Decision No One Has a Clean Answer To
When an enterprise team wants to deploy an AI agent to handle customer inquiries, reset passwords, or initiate payments, someone in that room needs to answer a specific question: is this use case safe to automate at this level of autonomy? Most governance teams do not yet have a repeatable method for answering it. They have AI ethics policies that speak to fairness and transparency, and general risk principles borrowed from software delivery, but neither maps directly to the question of how much autonomous action a given agent should be permitted to take.
That gap is consequential. An agent that only reads data and drafts a recommendation carries very different risk from one that reads customer records, calls an external API, and moves money. The same model, the same provider, and even the same code base can occupy either role, and the governance controls appropriate for one are inadequate for the other. Risk tiering is the upstream decision on which human oversight design, permission policies, and monitoring configuration all depend. For the broader operating model, see What Is AI Agent Governance.
The Gap in Most Enterprise Governance Approaches
General AI governance frameworks address that question in principle, but the practitioners making deployment decisions day-to-day need something they can apply to a specific use case, a specific autonomy level, and a specific set of connected systems.
The IMDA Model AI Governance Framework for Agentic AI (MGF for Agentic AI), published by Singapore's Infocomm Media Development Authority on 20 May 2026 and updated on 5 June 2026, is designed to fill exactly that space. Developed with feedback from over 60 organisations including AWS, DBS, Google, and Salesforce, the MGF is best-practice guidance (not a legally binding requirement) that gives organisations a structured way to assess and bound the risks of an agentic deployment before it reaches production.
The part of the MGF most relevant to deployment teams is its separate factor lists for impact severity and likelihood of failure. The MGF provides these as inputs to structure the assessment, not as a prescribed scoring matrix. One operational way to apply them is as a grid: placing a use case on the impact axis and the likelihood axis, then using the intersection to guide the starting risk level and governance intensity. The Dayos, OCBC, and MSD implementations in the MGF v1.5 each show a different way to do this in practice.
The Two Axes of Agentic AI Risk
Factors That Determine Severity of Impact
Impact severity concerns what happens downstream if the agent makes an error, acts on a manipulated input, or exceeds its intended scope. The MGF lists five factors that shape this dimension. The common thread across all of them is: how bad is the blast radius, and can it be contained?
Factor | Why it matters | Enterprise example |
|---|---|---|
Domain and use case | Tolerance for error in the domain and criticality of the business process the agent supports | A clinical documentation agent has lower error tolerance than a meeting scheduler |
Access to sensitive data | Whether the agent can reach personal, financial, or regulated data, especially with persistent memory across sessions | A customer-service agent with access to full account history |
Access to external systems | Whether the agent can act on external systems, creating potential for data leakage or disruption of third parties | An agent that writes to a partner database or sends external emails |
Scope of actions | Read versus write; a few pre-defined tools versus a broad action space | An HR agent that can update payroll across an employee group has broader scope than one that only reads records |
Reversibility of actions | Whether modifications can be easily undone, or trigger downstream obligations | External email delivery can be difficult or impossible to reverse once delivered, especially when it triggers downstream obligations |
Reversibility is often the most decision-relevant single factor in the impact column, because it determines the recovery options available after an error occurs. An agent that moves money, sends a message, or modifies a production record does something that may not be recoverable at all, or only through a costly process. That asymmetry between before and after execution is what elevates those use cases into higher governance tiers even when the probability of failure seems low. The full policy implications are covered in Reversibility Is a Governance Requirement.
Factors That Determine Likelihood of Risk Manifesting
Likelihood of failure is separate from impact severity, and the MGF for Agentic AI treats likelihood and impact as separate dimensions. The factors here concern the operational characteristics of the agent and its environment. They are not about what happens if things go wrong, but about how likely they are to go wrong.
Factor | Why it matters | Enterprise example |
|---|---|---|
Level of autonomy | Whether the agent follows a defined procedure or defines the entire workflow itself | A fully automated pipeline vs. one that routes every action for approval |
Task complexity | Number of steps and depth of analysis required at each step | Resolving a software defect versus resetting a password |
Access to external systems | Exposure to external, possibly untrusted, systems, increasing susceptibility to prompt injection and cyberattack | An agent with access to remotely hosted tools or external APIs vs. one querying an internal, read-only database |
Third-party involvement | Whether the agent relies on a vendor-operated service, limiting visibility and control | A vendor-embedded agent running inside a SaaS workflow |
System complexity | Multi-agent pipelines can produce emergent failures that individual agent assessments do not catch | An agent that delegates to sub-agents, each holding its own permissions |
The IMDA MGF for Agentic AI v1.5 (published 20 May 2026, updated 5 June 2026) added guidance on multi-agent systemic risk and third-party agent exposure, drawing on industry feedback from over 60 organisations. Both factors amplify every other likelihood dimension: a complex orchestration layer running an externally hosted sub-agent combines several risk vectors that may not be apparent from any individual agent assessment. For the specific problems this raises in pipeline governance, see Govern the Handoffs Between Your AI Agents.
Applying the Two Axes Together: The OCBC Source of Wealth Example
The OCBC Bank of Singapore case study in the MGF v1.5 demonstrates how this two-axis analysis shapes an agentic deployment decision in financial services. Bank of Singapore, a wholly owned private banking subsidiary of OCBC, developed an agentic AI system to support Relationship Managers and compliance teams in synthesising customer-provided financial documents into a structured Source of Wealth memo.
Working through the impact axis: the system handles sensitive financial data (income submissions, property documents, dividend statements), acts in a context where outputs inform onboarding and compliance decisions, and has limited reversibility once a Source of Wealth assessment influences downstream processes. Applying the MGF’s impact factors, an enterprise reviewing this use case might assess impact as high. Working through the likelihood axis: the pipeline uses narrowly scoped agents (each handles one function, from income document extraction to formatting), is triggered only by predefined workflows with no self-initiation, and includes a Human-in-the-Loop step at stage 3 of 8. Likelihood of uncontrolled failure might be assessed as moderate.
That composite picture points toward a governance design: the system provides decision support only and does not make credit, onboarding, or risk decisions autonomously. Final validation and approval remain with designated human reviewers at the end of the pipeline. The MGF case study describes this as task-level autonomy only: agents perform narrowly scoped tasks (extraction, drafting, checking), not end-to-end decisions. This is the intersection of high impact and moderate likelihood, resolved through scoped permissions and mandatory human sign-off, not through automation avoidance.
Translating Risk Scores into Deployment Tiers
The Dayos Three-Tier Model: Tiering Every Ticket by Risk Level
Dayos, an enterprise AI automation company headquartered in Singapore, contributed a case study to the IMDA MGF v1.5 (p. 18) showing how the two-axis scoring translates into a deployable tiering structure. Dayos built an AI-powered ticketing agent to handle every internal IT support request. Before any of this went live, every type of IT ticket was scored against three questions drawn from the MGF framework: severity of impact (if the agent gets this wrong, how bad is it?), reversibility (can the action be undone?), and feasibility of human oversight (is it realistic for a human to review this at each step?). For a broader framework connecting assessment to runtime enforcement, see the Governing AI Agents guide.
The scores determined which tier each ticket type falls into, and that tier dictated the reasoning strategy the agent uses and how much autonomy it gets.
Dayos Agentic Ticketing Tiers (IMDA MGF v1.5, p. 18)
Tier | Risk profile | Ticket examples (per MGF) | Agent behaviour | Oversight model | Share |
|---|---|---|---|---|---|
1 | Low severity, fully reversible | Password resets, access requests, status inquiries | Fully automated by agent using Simple Feedback (a propose-confirm loop). No human engineer in the loop. | Designated reviewer audits a cross-section of reasoning chains and confidence scores biweekly | 60% |
2 | Moderate severity, partially reversible | Chart of accounts updates, integration mapping corrections, diagnosing failed API connections | Agent diagnoses and proposes fix using ReACT (multi-step diagnostic loop). Human engineer must sign off before anything executes. | Qualified engineer approval before execution | 30% |
3 | High severity, limited reversibility | Production deployments, security changes, permission modifications | Agent does not touch these. Safeguards are expected to be validated before reassessment. | Out of scope for the agent at current maturity | 10% |
Excluding the 10% of tickets the agent never touches, 90% of total volume is handled entirely or in part by the agent. The MGF v1.5 case study (p. 18) documents that the full retirement of Dayos's own ServiceNow ITSM instance took 45 days and reduced legacy licensing costs by $121,000 annually.
Two design principles from this implementation carry forward. First, the scoring criteria were fixed before deployment: severity, reversibility, and oversight feasibility are properties of the ticket type, not outputs of the model. Second, for Tier 3, the case study documents that the agent does not touch these tickets at the current maturity of agentic AI. The MGF recommends more broadly that such boundaries be enforced through deterministic, system-level controls rather than prompt instructions alone, with access controls preventing prohibited tool calls rather than relying on the model to decline.
The MSD Agency Level Matrix: Five Levels of Agency Plus a Baseline
MSD, the pharmaceutical company that develops medicines, vaccines, and health solutions for 75,000 employees, contributed a second governance architecture to the MGF v1.5 (p. 21). Rather than a use-case-level tier model, MSD implemented an enterprise-wide Agentic System Level Matrix that maps five active levels of agency (L1 to L5), with L0 as the baseline for no agency, against seven governance dimensions, and uses an agent's level to determine the governance pathway intensity required.
MSD Agentic System Level Matrix (IMDA MGF v1.5, p. 21)
Dimension | L0 No Agency | L1 Assisted | L2 Supervised | L3 Conditional | L4 Advanced | L5 Full Agency |
|---|---|---|---|---|---|---|
Action Scope | N/A | Limited | Moderate | Moderate | Extensive | Extensive |
Permission Scope | N/A | Read-Only | Controlled Write | Controlled Write | Full Access | Full Access |
Autonomy | N/A | Reactive | Proactive | Proactive | Self-Directed | Self-Directed |
Goal Complexity | N/A | Simple | Simple | Compound | Compound | Dynamic |
Human Oversight | N/A | Continuous | Periodic | Exception-Only | Exception-Only | Exception-Only |
Risk Impact | N/A | Low | Medium | Medium | High | High |
Data & Context Scope | N/A | Task-Specific | Task-Specific | Cross-Application | Organisational | Organisational |
Lower levels (L0-L1) move through a lightweight governance pathway. Middle levels (L2-L3) move through an established AI impact assessment process. Higher levels (L4-L5) require escalations to enterprise architecture review and testing. The MGF case study notes that a programmatic runtime policy enforcement layer is being implemented at the AI gateway before higher levels of autonomy are enabled.
MSD also documented a containment strategy for third-party agent platforms. Rather than allowing vendor-embedded agents to act freely across MSD's data and systems, agentic features in third-party SaaS tools are by default restricted to within their own ecosystems. This contains residual risk while preserving the productivity benefits of approved use cases. The structural principle is the same as the MGF's general recommendation and the same as what the AI Agents Can Rewrite Their Own Permissions analysis shows: permission architecture, not prompt instructions, is the reliable control.
Designing Your Own Tiering System
The Dayos model and the MSD matrix represent two different operationalisations of the same underlying principle: classify by autonomy and impact, and enforce the classification through deterministic controls. The specific tiers, dimensions, and governance intensities need to reflect your organisation's risk appetite, regulatory requirements, and agent platform capabilities. The MGF does not prescribe a universal tier count or fixed scoring thresholds; it establishes the design logic.
Three design decisions every governance team needs to make:
What are the fixed criteria for tier assignment? They should be observable, stable properties of the action type (reversibility, data class, external system involved), not model-side metrics that can shift between runs.
At what level are tier constraints enforced? The MGF recommends preferring deterministic, system-level limits over prompt-based instructions alone: an access control that prevents the agent from calling a tool entirely is stronger than an instruction telling the agent not to use it.
Who is the named human for approval-gated operations, and what is the maximum response time? A gate that can be bypassed by waiting long enough is not a control.
Common Risk Factors That Organisations Underweight
Reversibility Is the Most Underrated Factor
In practice, reversibility has the most leverage of any single risk factor, because it determines the recovery options available after an error occurs. Calendar scheduling, password resets, and database reads are reversible or carry negligible downstream consequences. External email delivery, payment initiation, and record deletion vary: some can be partially reversed through additional process steps, but the time cost and downstream obligations involved often make recovery costly or impossible. A lower inherent risk score does not override low reversibility.
The MGF case study methodology reflects this consistently. Every tier boundary in the Dayos implementation references reversibility directly: fully reversible at Tier 1, partially reversible at Tier 2, limited reversibility at Tier 3. The criteria are properties of the action type, not performance metrics of the model executing it.
For governance teams designing their own frameworks, the practical starting point is a reversibility inventory: list every action class the agent can take, mark each as reversible, partially reversible, or irreversible, and treat irreversibility as an automatic escalation flag regardless of other impact scoring. The full implications for human-in-the-loop approval design follow directly from that classification.
Third-Party Agent Platforms Add a Hidden Risk Layer
When the agent being deployed relies on a third-party agent service, a vendor-embedded agent runtime, or an externally hosted model provider, its effective risk profile includes the risk surface of that third party. The IMDA MGF for Agentic AI v1.5, published 20 May 2026 and updated 5 June 2026, added guidance on this, identifying third-party agent usage as a likelihood-amplifying factor.
The MSD case study shows one concrete response: restrict vendor-embedded agents to their own ecosystems by default, rather than allowing cross-platform actions that may not have been exhaustively tested. This containment strategy keeps residual risk bounded while preserving productivity gains. For the underlying mechanism, see agent identity and trust propagation in the OpenBox docs.
System Complexity Can Create Compounding Failure Modes
Multi-agent pipelines can produce emergent and compounding failure modes that are not apparent from isolated agent assessments. A pipeline of three agents, each within individually acceptable risk parameters, can generate interactions that no single-agent review would flag. The MGF identifies agentic system complexity and multi-agent interactions as risk-amplifying factors, which is why the assessment needs to cover the pipeline, not each agent in isolation. For governing handoffs and trust delegation, see Govern the Handoffs Between Your AI Agents.
Residual Risk and Sign-Off Requirements
After controls are applied, some risk remains. A Tier 2 approval gate with a 30-minute response SLA still carries the residual risk of a delayed decision. A Tier 3 boundary that excludes certain ticket types still requires a human to handle the work the agent cannot do. Sound risk management expects organisations to evaluate residual risk explicitly and to document a named owner who accepts it before a deployment is approved. The OCBC case study illustrates this: even with decision support only and human final approval, the system operates under a named governance structure where outputs are explicitly advisory and subject to human validation.
Residual risk evaluation should answer three questions. First, what is the most likely failure mode that the controls in place would not catch? Second, what is the earliest observable signal that such a failure has occurred? Third, who is authorised to halt the agent if that signal appears?
These questions do not resolve automatically from the risk score. They require a named person, a documented signal, and a clear halt authority. An agent that processes work continuously does not wait for the next governance review cycle, so the halt authority needs to be exercisable promptly.
How OpenBox Operationalises the Risk Assessment Decision
An assessment recorded in a spreadsheet is a governance document. An assessment whose conclusions are enforced at runtime is a governance control. A common gap in enterprise agentic deployments is the separation between those two states.
OpenBox, an AI agent governance platform, is built to close it. The platform’s Trust Lifecycle begins with an Assess phase that evaluates an agent’s inherent risk and produces a Risk Profile Score. That score feeds the overall Trust Score formula: Trust Score = (Risk Profile Score × 40%) + (Behavioral Score × 35%) + (Alignment Score × 25%). The composite Trust Score determines the agent’s Trust Tier, which governs how strictly the platform enforces permissions and how much human oversight is required at runtime.
Risk assessment findings from the Assess phase connect directly to the Authorize phase, where Guardrails, stateless Policies (OPA/Rego), and stateful Behavioral Rules translate the tier classification into controls that operate on every agent action. At runtime, OpenBox returns one of five governance decisions: ALLOW, CONSTRAIN, REQUIRE_APPROVAL, BLOCK, or HALT, with precedence running HALT above BLOCK above REQUIRE_APPROVAL above CONSTRAIN above ALLOW. In an OpenBox deployment, a Tier 2 approval gate could be implemented as REQUIRE_APPROVAL and Tier 3 exclusions as BLOCK or HALT, depending on the control design. The Monitor phase then provides real-time observability, the Verify phase validates goal alignment through drift detection and attestation, and the Adapt phase evolves governance policies based on observed behaviour.
OpenBox supports and enforces what governance teams decide. The risk assessment judgment that determines which tier a use case belongs to, and what the residual risk owner accepts, remains with the humans accountable for the deployment.
Note on classification systems: the Dayos tiers classify action types by risk level for a specific use case; the MSD agency levels classify the overall autonomy of an agent deployment enterprise-wide; OpenBox Trust Tiers are a distinct, product-level classification derived from the composite Trust Score. These are complementary frameworks, not the same taxonomy.
Pre-Deployment Risk Assessment Checklist
The following example checklist draws on the MGF’s recommendations and extends them with suggested implementation controls. Items 1 through 8 correspond to practices the MGF recommends. Items 9 through 12 go further: residual-risk acceptance, halt authority, re-assessment triggers, and tamper-evident audit records are not prescribed as universal controls by the MGF but represent common enterprise practice. Adapt all items to your organisation’s regulatory context and risk appetite.
# | Pre-deployment check |
|---|---|
1 | The use case has been scored on both impact severity and likelihood of failure as separate dimensions, not combined into a single assessment. |
2 | Reversibility has been assessed for every action class the agent can take, and irreversible actions are explicitly identified. |
3 | Sensitive data access has been documented: what data categories the agent can read, what it can write, and which records it cannot touch. |
4 | External system access has been mapped: every API, tool, MCP server, and downstream system the agent can reach, including third-party services. |
5 | A named tier assignment has been made for this use case, with documented criteria for that assignment. |
6 | The Tier 2 approval gate (where applicable) has a named human approver and a documented maximum response time. |
7 | The Tier 3 boundary (actions excluded from the agent's scope) is enforced at the system or permission level, not only in the agent prompt. |
8 | Third-party or vendor-embedded components have been assessed separately and a containment plan is in place if they behave unexpectedly. |
9 | A named owner has accepted the residual risk that remains after controls are applied. |
10 | A halt authority and trigger condition have been documented: who can stop the agent and what observable signal activates that decision. |
11 | The first re-assessment trigger has been set (capability change, data access expansion, or following any incident). |
12 | A tamper-evident audit record will be kept of the agent's runtime decisions for the duration of operation, with defined retention. |
Frequently Asked Questions
What makes agentic AI risk assessment different from standard software risk assessment?
Standard software risk assessment covers system outputs, failure modes, and system-level risks such as data integrity, availability, and security vulnerabilities. Agentic AI risk assessment adds a distinct dimension: the system's ability to take autonomous actions, invoke tools, alter external state, and operate without step-by-step human oversight. An agent that can call tools, execute workflows, and take multi-step actions can cause harm before any human reviews the result. That means the assessment must cover actions and their reversibility, not only model outputs and their accuracy. The two types of failure are fundamentally different in their timing, detectability, and corrective options.
How do I score reversibility for agent actions that have downstream dependencies or contractual implications?
Treat an action as irreversible if either of two conditions applies: the action itself cannot be undone, or undoing it would require a process that is slower or more costly than the original harm. A payment can sometimes be reversed, but typically only through a dispute process that takes days and may involve a third party. For risk assessment purposes, classify it as irreversible and design the governance controls accordingly. Downstream contractual implications, such as an email that triggers a commitment, amplify the irreversibility even when the underlying technical operation could theoretically be cancelled.
What is the practical difference between impact assessment and likelihood assessment in the two-axis framework?
Impact assessment asks: what is the worst credible outcome if this agent acts incorrectly, and how far does the harm extend? Likelihood assessment asks: given the agent's autonomy level, task complexity, and system dependencies, how probable is that failure? The IMDA MGF treats the two as separate dimensions. A low-probability, catastrophic-impact event and a high-probability, minor-impact event may both be acceptable at different risk tolerances, but for entirely different reasons. Conflating them into a single score produces a governance decision that cannot be explained or defended if the deployment is later challenged.
What is residual risk, and who should sign off on it?
Residual risk is the exposure that remains after all controls, tier boundaries, approval gates, and monitoring have been applied. It is rarely zero. For a Tier 1 automated ticketing agent, the residual risk might be a correctly classified low-severity ticket that carries an unusual downstream dependency the classification logic did not capture. The sign-off on residual risk should come from the person with operational accountability for the function the agent serves, along with security or compliance sign-off where regulated processes are involved. The assessment team should not be the only party accepting the residual.
Why should third-party-operated agents be assessed separately from internally developed ones?
When an agent relies on a third-party agent service or vendor-embedded agent runtime, its effective risk profile includes the risk surface of that external service. The governance question is no longer only about what the internal agent will do, but also about what happens when the external service behaves unexpectedly, returns an unexpected result, or becomes unavailable. Containment design for third-party agent integrations is a distinct sub-problem from permission design for the agent itself. The IMDA MGF for Agentic AI v1.5, published 20 May 2026 and updated 5 June 2026, added guidance specifically addressing third-party agent usage as a likelihood-amplifying factor.
Sources |
IMDA, "Updated Model AI Governance Framework for Agentic AI" (v1.5), https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/factsheets/2026/updated-model-ai-governance-framework-for-agentic-ai, accessed 21 September 2026. |
MDDI (Singapore), "Singapore Launches New Model AI Governance Framework for Agentic AI", https://www.mddi.gov.sg/newsroom/singapore-launches-new-model-ai-governance-framework-for-agentic-ai--/, accessed 21 September 2026. |
OpenBox (docs.openbox.ai), "Governance Decisions", https://docs.openbox.ai/core-concepts/governance-decisions, accessed 21 September 2026. |
OpenBox (docs.openbox.ai), "Trust Scores", https://docs.openbox.ai/core-concepts/trust-scores, accessed 21 September 2026. |
OpenBox (docs.openbox.ai), "Trust Tiers", https://docs.openbox.ai/core-concepts/trust-tiers, accessed 21 September 2026. |
OpenBox (docs.openbox.ai), "Trust Lifecycle", https://docs.openbox.ai/trust-lifecycle, accessed 21 September 2026. |
OpenBox (docs.openbox.ai), "Assess (Trust Lifecycle Phase 1)", https://docs.openbox.ai/trust-lifecycle/assess, accessed 21 September 2026. |

