AI Governance & Compliance
What Your LLM Gateway Can’t See: Agent Tool Governance
An approved model call can still trigger unapproved tool actions. Here is what lives outside the LLM gateway boundary and how to govern it.
Published on


The LLM Gateway Can't See the Real Attack Surface: Governing the Tools Around Your Model Calls
A provider allowlist enforced at the gateway solves the routing problem. It does not solve what happens in the execution path that follows the model call.
By Tahir Mahmood, Co-founder & CTO, OpenBox · Last updated 15 September 2026
A compliant model call and an unapproved side effect can happen in the same run
Suppose your OpenRouter agent is configured with a provider allowlist. The allowlist restricts routing to the approved providers; to make that restriction fail-closed, `allow_fallbacks` must also be disabled. From a routing-governance perspective, these are the correct controls.
Now suppose the same agent has a tool that posts to an external HTTP endpoint. The model receives the prompt, decides to call that tool, and the tool sends a payload to a URL that nobody on your team reviewed.
The routing record shows a clean run. The provider was approved. The gateway honored the allowlist. None of that is inaccurate.
The data went somewhere your organisation did not authorise. The gateway has no record of it, because the tool's HTTP request did not traverse the gateway.
This is the structural issue that gateway-only governance leaves open. A provider allowlist controls one class of routing risk. It does not constrain what agent tools do with the model output they receive. Governing the model provider and governing the agent are related but technically separate problems.
What an LLM gateway actually governs
An LLM gateway operates at the boundary between the agent and the model provider. OpenRouter receives model requests, routes them upstream, returns the response, and publishes a generation record per call containing the provider that served it, the data region, the cost, and any fallback attempts. This is documented and verifiable evidence.
What the gateway does not see is the activity inside the agent process after the model responds. The OpenBox Proof of Routing documentation describes its own scope precisely: it adds a provenance layer over “the one decision those cannot see: which upstream provider a gateway picked, after the request left your process.”
When the agent runtime invokes a tool following a model response, that tool execution happens in the calling process. An outbound HTTP call the tool makes, a database query it issues, a file it reads: none of these cross the model-routing gateway. A model-routing gateway sits on the path to the model provider. It does not inherently govern every path the agent process can take.
This is a boundary, not a failure. The gateway does exactly what it was designed to do. The point is that its design scope does not extend to instrumenting arbitrary activity in the process that calls it.
The execution surfaces that live outside the gateway boundary
An agent built on the OpenRouter Agent SDK can act across several surfaces in a single run: model calls, tool invocations, and from within those tools, outbound HTTP requests, database operations, and file activity. Each can carry information derived from, or otherwise influenced by, the model’s inputs or outputs. The OpenBox OpenRouter SDK Reference documents the surfaces the governance layer covers in a governed session:
Execution surface | Crosses model gateway? | Governed by OpenBox OpenRouter SDK |
|---|---|---|
Model calls | Yes | Yes, via openbox.callModel() |
Tools | No | Yes, via openbox.tools() |
HTTP requests | No | Yes, when inside a governed operation |
Database operations | No | Yes: pg, mysql2, mongodb, redis, ioredis |
File I/O | No | Optional; off by default |
The surfaces that do not cross the model gateway are where an agent's external side effects can occur. An execution path that illustrates the gap: the agent sends a prompt to an approved provider (the gateway records this); the model responds with a tool invocation; the tool posts to an endpoint not on any approved list. Steps two and three are invisible to the routing gateway. The gateway's record is accurate for what it covers. The risk lives in what it does not cover.
Watching a tool and controlling a tool are different things
There is a meaningful difference between recording that a tool ran and deciding whether it should run before it does.
openbox.tools() wraps each tool so that its execution is evaluated before it runs, the call can be held open across a human approval, and the execution is recorded when it finishes. Per the OpenBox documentation on getting started with OpenRouter, this is the end-to-end governance model: evaluated before, approvable during, recorded after.
When a governance policy returns REQUIRE_APPROVAL for a tool operation, the call is held open. The tool body does not execute until the approval result arrives. A human reviews the pending operation in the dashboard and approves or rejects it; on rejection, the operation is stopped. A model-routing gateway, by itself, does not provide this control. The tool execution happens in the calling process, not at a network boundary the gateway can intercept. A system that captures tool output after the tool runs records what happened. A governance layer that holds the tool before it runs controls what happens. These are different properties.
OpenBox applies five governance decisions to tool operations and every other governed operation. In precedence order: HALT terminates the agent session; BLOCK stops the specific operation; REQUIRE_APPROVAL pauses it for human review; ALLOW permits it to proceed. A fifth verdict, CONSTRAIN, runs the operation inside a sandbox rather than stopping it outright; it sits between ALLOW and BLOCK in the precedence chain. A tool call can receive any of these decisions before its body executes.
One attested session record across multiple execution surfaces
The governance property that matters most here is that model calls and tool operations can belong to the same session, be evaluated under the same policy, and produce a single attested record.
When openbox.callModel() and openbox.tools() both operate in a single run, the session record spans the model call, the routing provenance from OpenRouter's generation record, and the tool executions that followed. These are not separate telemetry streams in separate systems; they belong to one governed session context.
This distinction matters for accountability. A routing record answers where the model call went. A tool execution record answers what the tool did. Establishing the causal chain, that this model response from this provider led to this tool invocation which sent data to this endpoint, requires a session record that spans all three.
OpenBox seals this session record using cryptographic attestation. Per the documentation, each event is hashed with SHA-256, combined into a Merkle tree, and the session root is signed using ECDSA NIST P-256 via AWS KMS by default. The signed root makes post-session alteration of the attested record detectable after the session closes.
The routing provenance from OpenRouter's generation record, which the SDK reads and hashes into the same Merkle tree, is part of that sealed context, not a separate record held by a separate system.
Two questions govern an agent, not one
Two questions arise when an agent runs, and they are related but technically distinct.
The first: which model provider did this prompt reach? A provider allowlist, Proof of Routing, and OpenRouter's generation record answer this. The evidence is re-checkable at the gateway.
The second: given that model response, what did the agent do next, and was it authorised? Answering this requires governance that follows the execution graph beyond the model call, covering the tools the model invokes and the operations those tools perform.
A gateway governs the edge from the agent to the model provider. Session-level governance covers the execution path that follows.
The OpenBox OpenRouter integration, covered in detail in the OpenBox integration guide, wraps both: model calls via openbox.callModel() and tools via openbox.tools(). OpenRouter remains the routing boundary; OpenBox, an AI agent governance platform, adds policy evaluation, human approval gates, and attested recording across the broader execution surface.
The case for doing both is not that gateways are insufficient for their purpose. It is that their purpose is narrower than governing the full execution path of an agent.
Frequently Asked Questions
Does a provider allowlist prevent tool-based data exposure?
A provider allowlist controls where model calls are sent; it does not constrain what agent tools do with the model output afterward. When a tool makes an outbound HTTP call or queries a database, that activity happens in the calling process, outside the gateway's scope. Tool execution therefore needs controls beyond the model-routing gateway when those actions fall within the system's risk boundary.
Can REQUIRE_APPROVAL pause a tool before its body executes?
Yes. When openbox.tools() wraps a tool and a policy returns REQUIRE_APPROVAL, the tool call is held open while a human decides; the tool body does not execute until the approval result is received. Two tool shapes, an async-generator execute and a manual tool with no body, use hook-based governance instead and cannot be held open across an approval.
What execution surfaces does the OpenBox OpenRouter SDK instrument?
Per the OpenBox OpenRouter SDK Reference, the governed surfaces are model calls, tools, HTTP requests inside a governed operation, the database drivers pg, mysql2, mongodb, redis, and ioredis, and file I/O, which is off by default. These surfaces are what an agent uses to move information and can each carry what the model received.
Where does routing provenance sit relative to the tool execution record?
The OpenBox SDK reads OpenRouter's Generation API for each governed model call and records the result as span attributes, included in the same session Merkle tree as the tool execution events. The session root is signed at session close, so routing provenance forms part of the same attested session context as tool and HTTP activity records.
Sources |
OpenBox (docs.openbox.ai), "Getting Started with OpenRouter," https://docs.openbox.ai/getting-started/openrouter, accessed 15 September 2026. |
OpenBox (docs.openbox.ai), "OpenRouter SDK Reference," https://docs.openbox.ai/developer-guide/openrouter/sdk-reference, accessed 15 September 2026. |
OpenBox (docs.openbox.ai), "Proof of Routing," https://docs.openbox.ai/developer-guide/openrouter/proof-of-routing, accessed 15 September 2026. |
OpenBox (docs.openbox.ai), "Routing Policies," https://docs.openbox.ai/developer-guide/openrouter/routing-policies, accessed 15 September 2026. |
OpenBox (docs.openbox.ai), "Governance Decisions," https://docs.openbox.ai/core-concepts/governance-decisions, accessed 15 September 2026. |
OpenBox (docs.openbox.ai), "Attestation & Cryptographic Proof," https://docs.openbox.ai/administration/attestation-and-cryptographic-proof, accessed 15 September 2026. |
OpenBox (docs.openbox.ai), "Guardrails," https://docs.openbox.ai/trust-lifecycle/authorize/guardrails, accessed 15 September 2026. |

