AI Governance & Compliance

Structural Guardrails for Enterprise Agentic AI

Prompt instructions don't reliably constrain AI agents at scale. Learn why deterministic, structural controls, and MCP as a governance enforcement layer, are the enterprise standard.

Published on

Subscribe to our newsletter

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Prompt-Layer Guardrails Are Not Enough: Building Structural Controls for Enterprise Agentic AI

System prompt instructions are a common safety layer for AI agents in production. They are also probabilistic in their enforcement, inconsistent across deployments, and degradable under adversarial pressure. Singapore’s IMDA Model AI Governance Framework for Agentic AI now makes the distinction explicit: structural, deterministic controls belong at the infrastructure layer, and prompt-layer safeguards are one input among several, not the primary enforcement mechanism.

By Tahir Mahmood, Co-founder & CTO, OpenBox  ·  Last updated 3 October 2026

The Guardrail Illusion: What Prompt Instructions Cannot Guarantee

A large share of production AI agents rely on system prompt instructions as their primary safety mechanism. A system prompt tells the agent what it should and should not do: do not access customer financial data, always request approval before sending emails, never execute shell commands without user confirmation. These instructions work, until they do not.

Prompt-layer safeguards produce probabilistic enforcement. The prompt itself is a static artifact, but the model’s adherence to it is not deterministic. The same instruction can produce different compliance outcomes across model versions, temperature settings, context window lengths, and prompt structures. Under adversarial input, including prompt injection, jailbreaking, and instruction-following conflicts, compliance degrades further.

The IMDA Model AI Governance Framework for Agentic AI (v1.5), published 20 May 2026 and updated 5 June 2026, addresses this directly in Section 2.3.1. The framework recommends that organisations “rather than prompt-layer safeguards, consider implementing deterministic safeguards that operate at a system-level through predefined logic, especially for higher-risk actions.” It notes that prompt-layer safeguards “tend to be inconsistently defined across users, as opposed to system-level safeguards that can be consistently defined and enforced.”

This is not a theoretical concern. When 50 engineering teams define 50 versions of the same behavioural constraint in 50 different system prompts, no central governance record exists. A compliance audit cannot reconstruct which version of which constraint was active for which agent at which time. The prompt was the control, and the prompt was never versioned, never signed, and never enforced independently of the model that received it.

When Prompt-Layer Controls Are the Right Choice

None of this means prompt-layer safeguards have no value. They remain appropriate for nuanced content moderation, open-ended behavioural requirements that resist rule-based encoding, and defence-in-depth layering where structural controls handle the hard boundaries and prompt-layer instructions handle softer guidance. The MGF distinguishes between control types precisely so that practitioners can match the right mechanism to the right risk, not to eliminate one category entirely.

Agentic AI Technical Controls: The Three-Layer Architecture

The MGF (Section 2.3.1) distinguishes between structural and rule-based controls, model-based or prompt-layer controls, and runtime controls. This three-layer taxonomy provides a useful organising principle for agentic AI technical controls. Each layer operates differently, fails differently, and is appropriate for different risk profiles.

Layer 1: Structural, Rule-Based Controls (Deterministic)

Structural controls enforce policy through predefined logic that operates independently of the agent’s reasoning. An access-control rule that prevents an agent from calling a database write API is structural: it does not depend on the model understanding or obeying an instruction. The agent cannot override it by reasoning differently, by receiving adversarial input, or by updating its own prompt.

The enforcement logic in these controls is deterministic: it produces the same decision regardless of which model version, prompt configuration, or context window is in use, even though the agent’s attempted actions may vary. Whether these controls are also auditable depends on implementation: proper instrumentation, logging, and evidence preservation are required to make any control auditable in practice. Implementation typically occurs at the tool layer, the API gateway, or the protocol layer.

The MGF’s stated preference for higher-risk actions is clear: the framework recommends considering structural controls over prompt-layer instructions. Rather than relying on a system prompt to instruct an agent not to access certain tools, the framework recommends imposing access controls that prevent the tool from being called by the agent at all.

Layer 2: Model-Based and Prompt-Layer Controls (Probabilistic)

Model-based controls operate through the agent’s language model, either as system prompt instructions or as a separate model-based judge that evaluates actions before or after execution. These controls are flexible and can handle nuanced, context-dependent requirements that resist deterministic encoding.

Their limitation is probabilistic enforcement. A prompt-layer instruction shapes model behaviour, but the model’s compliance is not guaranteed by the infrastructure. The model may follow the instruction with high reliability under normal conditions, but that reliability degrades under adversarial pressure, across model updates, and at the edges of the instruction’s intended scope. For enterprise governance, the critical gap is independent enforcement: prompt configurations can be version-controlled, tested, and reviewed, but whether the model actually complied on a given invocation cannot be independently verified at the infrastructure layer.

This makes prompt-layer controls appropriate as one layer in a defence-in-depth strategy, not as the sole enforcement mechanism for high-risk actions.

Layer 3: Runtime Controls (Dynamic)

Runtime controls operate during agent execution: monitoring, observing, and intervening in real time. They include rate limiting, input validation, anomaly detection, and automated escalation. Unlike static structural controls, runtime controls are adaptive. They can detect patterns that emerge only during execution, such as an agent making an unusual number of API calls, accessing data outside its normal scope, or exhibiting goal drift across a session.

Runtime controls complement, rather than replace, static structural controls. A structural control prevents an agent from calling an unauthorised API endpoint. A runtime control detects when an agent’s pattern of authorised API calls becomes anomalous and escalates for review. Both are necessary; neither is sufficient alone.

A Reference Control Catalog by Agent Component

The MGF (Section 2.3.1) provides sample controls organised by agent component: planning, tools, and protocols. The following table maps each component to the specific recommended control types, adapted for practitioner use. Note that the MGF presents these as illustrative controls, not an exhaustive or mandatory list.

Agent Component

Control Type

Implementation Approach

Risk Addressed

Planning

Prompt-Layer

Prompt agent to reflect on plan adherence; log plan and reasoning in structured format for human review

Erroneous or misaligned execution plans

Planning

Prompt-Layer

Prompt the agent to summarise its understanding and request clarification from the user before proceeding

Cascading errors from misunderstood goals

Tools

Structural

Configure strict validated input formats; enforce least-privilege access through authentication and authorisation at the tool layer

Unauthorised data access; tool misuse

Tools

Structural

Restrict write access to sensitive databases; require human takeover for sensitive data input

Data corruption; credential exposure

Protocols (MCP)

Structural

Whitelist trusted MCP servers; filter sensitive data at the MCP layer; log all agent-to-system interactions

Data exfiltration via untrusted servers

Protocols (MCP)

Structural

Sandbox code execution on MCP servers; enforce first-use trust verification for new servers

Arbitrary code execution; supply-chain compromise

Multi-Agent

Structural + Runtime

Use typed function calls and structured schemas (not free-text inter-agent instruction); implement inter-agent authentication

Cascading failures; instruction injection across agents

A key insight from this catalog: validating intent before execution is more efficient than detecting errors after actions have been taken. Logging an agent’s plan and requiring structured confirmation before high-risk operations shifts the governance boundary upstream, where intervention is cheaper and consequences are reversible.

MCP as a Governance Enforcement Layer

The Model Context Protocol (MCP), originally created by Anthropic and donated to the Agentic AI Foundation (AAIF), a directed fund under the Linux Foundation, standardises how AI agents discover and invoke external tools. Technical direction remains with the existing maintainer community; the Linux Foundation provides the neutral institutional home and infrastructure. MCP is primarily a connectivity protocol. The MGF, however, identifies a governance opportunity in its architectural position.

The MGF notes that MCP “can potentially act as a governance layer as it sits between the agent and the enterprise systems it accesses.” For agents that use MCP to connect to tools, controls defined at that layer (server whitelisting, data filtering, interaction logging, code sandboxing) apply consistently across all agents connecting through those servers.

Two qualifications are important. First, MCP is optional: not all agent tool calls traverse MCP, and agents can invoke tools through direct API calls, SDKs, or other protocols. The governance value applies only to the execution paths that run through MCP. Second, the MCP specification itself notes that “MCP itself cannot enforce these security principles at the protocol level”; enforcement is an implementation responsibility, not a protocol guarantee.

This matters for enterprise governance because, where agents do use MCP, it provides a natural enforcement point. Defining access policies, data-filtering rules, and audit requirements at an MCP gateway can enforce them for agents connecting through that gateway, without requiring equivalent constraints in each agent’s system prompt.

Structural Controls in Practice: Two Case Studies

The MGF v1.5 includes case studies from organisations that have implemented structural controls in agentic AI systems. Two are particularly instructive for the distinction between prompt-layer and system-level enforcement.

Tencent CodeBuddy: Action-Type Checkpoint Enforcement

Tencent’s CodeBuddy is an agentic AI coding system developed by Tencent Cloud, used by its engineers. CodeBuddy can plan, write, and deploy code through natural-language instructions and can access filesystems, terminal commands, external APIs, and MCP tools.

CodeBuddy implements structural approval requirements by action type, not by prompt instruction. The system defines by default which actions require human approval: editing files, running shell commands, making network requests, and using external tools all require explicit approval. Read operations, such as viewing files and listing directories, can proceed without approval under default permissions within trusted directories.

The design addresses a specific failure mode of prompt-layer controls: context-dependent bypass. Even when a general class of commands has been pre-approved for a session, if a suspicious variant appears (the MGF example: a command attempting to write to an unexpected file path), CodeBuddy’s structural protections require fresh human approval. The approval checkpoint is enforced at the system level, not left to the model’s interpretation of what constitutes “suspicious.”

CodeBuddy also implements plain-language explanation of complex commands before approval, so that the human reviewer can make a genuinely informed decision rather than rubber-stamping technical output they cannot evaluate quickly.

Terminal 3: Hardware-Level Enforcement for Financial Agents

Terminal 3, a Hong Kong-based data privacy infrastructure provider, deploys a finance payroll agent that automates monthly payroll cycles: computing salaries, validating expense claims, and executing financial transactions. The case study, published in the MGF v1.5 (Section 2.3.1), demonstrates hardware-level structural controls for high-stakes financial operations.

Terminal 3 implements four structural controls. First, a Verifiable Credential of Intent: before each payroll cycle, the HR director issues a cycle-scoped credential to the agent, defining exactly what the agent is authorised to do. Second, sensitive data (employee records, salary figures) is held exclusively in a Trusted Execution Environment (TEE) and never transmitted to the agent context, which architecturally eliminates prompt-injection-based data exfiltration for that data. Third, every action is verified against credentialed parameters at the hardware level. Fourth, each step of execution, including policy evaluations, is recorded in a cryptographically verifiable log to provide the evidence chain for post-incident review.

This architecture represents the far end of the structural-control spectrum. The agent’s reasoning is architecturally separated from execution. None of these controls depend on the agent’s prompt correctly instructing it to handle sensitive data carefully. The governance boundary is enforced by infrastructure, not by instruction.

Change Management for Complex Agent Systems

Structural controls solve the enforcement problem. They introduce a different problem: change management. The MGF (Section 2.3.3) addresses this directly, because changes in agent configuration can cascade in ways that traditional software change management may not anticipate. Traditional change management already accounts for unintended effects, but agentic systems introduce additional sources of behavioural variability: model-dependent output formatting, emergent tool-calling sequences, and dynamic interactions that are difficult to enumerate in advance.

A model update that changes how the agent formats tool call arguments can break downstream workflow integrations. A prompt refinement that shifts how the agent interprets authorisation boundaries can inadvertently enable actions that were previously constrained. A new MCP connection can introduce an untrusted data source into an otherwise controlled environment.

Defining Change Review Triggers

Effective change management for agentic systems requires defining what constitutes a reviewable change. Four categories of triggers apply:

Technical triggers: model updates, tool modifications, new MCP connections, protocol version changes, SDK updates, and changes to agent configuration files.

Environmental triggers: changes in the domain the agent operates in, business context shifts, and changes in the data sources the agent accesses.

Performance triggers: anomalous agent behaviour detected by runtime monitoring, degraded accuracy or goal alignment, and increased error rates in tool interactions.

Regulatory triggers: changes in compliance requirements, new regulatory guidance affecting the agent’s operational scope, and updates to applicable standards or frameworks.

Categorising Changes by Risk Level

Not all changes warrant the same review depth. The MGF’s risk assessment framework (Section 2.1) provides the basis for a tiered approach. The following is a practitioner model for categorising changes; it is not prescribed by the MGF.

Minor changes (prompt phrasing adjustments, non-functional configuration updates): lightweight review, documented but typically not requiring full governance sign-off.

Material changes (model version updates, tool additions, changes to the agent’s autonomy scope): consider a full governance review with risk assessment.

Critical changes (high-stakes workflow modifications, external system access expansion, changes to approval requirements): consider immediate re-assessment using the framework’s risk factors (Section 2.1.1), with sign-off from the relevant governance authority.

How OpenBox Operationalises Structural Controls

OpenBox, an AI agent governance platform, translates governance requirements into auditable, runtime-enforced policy, not prompt instructions. Two phases of the OpenBox Trust Lifecycle are directly relevant.

The Authorize phase is the enforcement layer. It configures guardrails (hard constraints on agent actions), policies (OPA/Rego stateless permission checks), and Behavioral Rules (stateful, multi-step pattern detection). These controls evaluate governed operations before execution and return one of five governance decisions: ALLOW, CONSTRAIN, REQUIRE_APPROVAL, BLOCK, or HALT, with precedence HALT > BLOCK > REQUIRE_APPROVAL > CONSTRAIN > ALLOW.

CONSTRAIN runs an action inside a sandbox rather than stopping it outright, sitting between ALLOW (which permits the operation without restriction) and REQUIRE_APPROVAL (which pauses the operation for human review). This gives governance teams a graduated response: an action too sensitive to run unrestricted but not risky enough to pause for human sign-off can execute within a controlled boundary. The enforcement is structural: it operates independently of the agent’s prompt and cannot be bypassed by the model’s reasoning.

The Verify phase provides the evidence layer. It validates that agents acted in alignment with their stated goals and produces cryptographic attestation for tamper-evident audit trails. Each session’s events are hashed into a Merkle tree and digitally signed, creating a verifiable proof certificate that governance data was not altered and that governance decisions were recorded accurately. Together, Authorize and Verify close the gap between governance intent and operational evidence: the structural control runs on every governed operation, and the cryptographic record provides tamper-evident evidence that it did.

Developer Implementation Checklist: Agentic AI Technical Controls

The following checklist provides a recommended set of structural controls for developers building or reviewing enterprise agentic AI systems. It draws on the MGF’s recommendations and the case studies above, but represents a practitioner synthesis, not an MGF-prescribed minimum.

#

Control

Layer

1

Enforce least-privilege tool access through authentication and authorisation at the tool or API layer, not through prompt instructions

Structural

2

Restrict agent write access to sensitive databases unless strictly required; default to read-only

Structural

3

Require human approval for high-stakes and irreversible actions, enforced at the system level with structured approval requests

Structural

4

Whitelist trusted MCP servers; block connections to unverified servers by default

Structural

5

Sandbox code execution on MCP servers and in agent environments executing untrusted input

Structural

6

Log agent plans and reasoning traces in structured format before execution of multi-step workflows

Structural

7

Implement inter-agent authentication and typed function calls for multi-agent interactions (no free-text inter-agent instructions)

Structural

8

Configure runtime anomaly detection for agent behaviour: unusual API call patterns, scope violations, and goal drift

Runtime

9

Define change review triggers for model updates, MCP connection changes, tool modifications, regulatory changes, and autonomy scope changes

Process

10

Categorise agent configuration changes by risk level (minor, material, critical) with corresponding review depth

Process

11

Maintain prompt-layer controls as a defence-in-depth complement for nuanced behavioural requirements, not as the primary enforcement mechanism for high-risk actions

Prompt-Layer

12

Establish cryptographic audit trails that record governance decisions with tamper-evident evidence, independent of the agent’s own logging

Structural

Control Layer Matrix: Matching Action Types to Controls

The following matrix maps common agent action types to a recommended control layer, based on the MGF’s risk factors (Section 2.1.1) and the three-layer taxonomy above. This is an illustrative practitioner model; actual risk levels depend on domain, data sensitivity, reversibility, and the specific deployment context.

Action Type

Typical Risk Level

Recommended Control Layer

Example Implementation

Read internal data

Low (varies by data sensitivity)

Structural (least-privilege access)

Scope API keys to read-only; log access

Write to database

Medium-High

Structural (access control + approval gate)

Require human approval; restrict write scope

Execute shell command

High

Structural (sandboxed execution + approval)

Sandbox environment; require per-command approval

Send external communication

High (irreversible)

Structural (approval gate)

System-level approval before send; log content

Financial transaction

Critical

Structural (credential-scoped + TEE)

Verifiable Credential of Intent; hardware execution boundary

Content generation

Varies (low for internal drafts; high for regulated or external-facing content)

Prompt-Layer + Runtime monitoring

Content guidelines in prompt; runtime quality checks

Multi-agent delegation

Medium-High

Structural (typed schemas + authentication)

Structured handoff; per-agent trust verification

The general pattern: as risk increases and reversibility decreases, the appropriate control layer tends to shift from prompt-layer guidance toward structural, infrastructure-level enforcement. This is a design heuristic consistent with the MGF’s recommendations, not a formal framework rule mapping each risk level to a specific control type.

Frequently Asked Questions

What is the difference between structural controls and prompt-layer controls for AI agents?

Structural controls enforce policy through infrastructure: access-control rules, API-layer restrictions, and protocol-level configurations that operate independently of the agent’s language model. Prompt-layer controls enforce policy through natural-language instructions in the agent’s system prompt. Structural controls produce deterministic enforcement decisions; prompt-layer controls depend on the model complying with the instruction, which is not guaranteed. The IMDA MGF for Agentic AI recommends considering structural controls for higher-risk actions.

Can MCP function as a governance enforcement point, or is it only a connectivity protocol?

MCP is primarily a connectivity protocol that standardises how agents interact with tools. The IMDA MGF identifies a governance opportunity in its architectural position: for agents that use MCP, controls defined at the MCP layer (server whitelisting, data filtering, interaction logging, code sandboxing) apply consistently across all agents connecting through those servers. However, MCP is optional (not all agent tool calls traverse it), and the MCP specification notes that security enforcement is an implementation responsibility, not a protocol-level guarantee. MCP is an industry protocol, not an OpenBox product.

How does change management differ for agentic AI systems compared to traditional software?

Traditional software change management already accounts for unintended effects, but agentic systems introduce additional sources of behavioural variability. A model update can alter tool call formatting, breaking downstream integrations; a prompt refinement can shift authorisation interpretation, enabling previously constrained actions; a new MCP connection can introduce untrusted data sources. Agentic change management benefits from defining specific triggers (technical, environmental, performance, regulatory) and categorising changes by risk level for appropriate review depth.

Should prompt-layer controls be eliminated entirely in favour of structural controls?

No. The MGF distinguishes between control types so that practitioners can match the right mechanism to the right risk. Prompt-layer controls remain valuable for nuanced content moderation, context-dependent behavioural requirements, and defence-in-depth layering. The recommendation is that prompt-layer controls should not be the sole or primary enforcement mechanism for high-risk actions. They are one layer in a multi-layer control architecture, not the foundation.

Sources

IMDA and AI Verify Foundation, “Model AI Governance Framework for Agentic AI,” Version 1.5, published 20 May 2026, updated 5 June 2026, https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf, accessed 3 October 2026.

MCP Blog, “MCP joins the Agentic AI Foundation,” 9 December 2025, https://blog.modelcontextprotocol.io/posts/2025-12-09-mcp-joins-agentic-ai-foundation/, accessed 3 October 2026.

MCP Specification (2026-07-28), https://modelcontextprotocol.io/specification/2026-07-28, accessed 3 October 2026.

OpenBox (docs.openbox.ai), “Governance Decisions,” https://docs.openbox.ai/core-concepts/governance-decisions, accessed 3 October 2026.

OpenBox (docs.openbox.ai), “Authorize Phase,” https://docs.openbox.ai/trust-lifecycle/authorize, accessed 3 October 2026.

OpenBox (docs.openbox.ai), “Verify Phase,” https://docs.openbox.ai/trust-lifecycle/verify, accessed 3 October 2026.

OpenBox (docs.openbox.ai), “Attestation & Cryptographic Proof,” https://docs.openbox.ai/administration/attestation-and-cryptographic-proof, accessed 3 October 2026.

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us