AI Governance & Compliance
AI Agent Lifecycle Management: A Production Checklist
AI agent lifecycle management with owners, gates and evidence, from registration to retirement. By OpenBox, the AI agent governance platform.
Published on


AI Agent Lifecycle Management: A Production Checklist
By Tahir Mahmood, Co-founder & CTO, OpenBox · Last updated 21 September 2026
AI agent lifecycle management is the process of registering, assessing, building, testing, approving, operating, changing and retiring an AI agent. Each stage should have an accountable owner, explicit action boundaries and evidence that the relevant controls work.
What is AI agent lifecycle management?
AI agent lifecycle management is a per-agent operating procedure. It records who owns an agent, what it may do, how those limits were tested, who approved the release, what changed afterwards and how the agent is switched off. The eight stages here are an OpenBox editorial model, not a NIST, IBM or AWS standard.
IBM's explainer, published on 23 June 2026, defines agent lifecycle management as the end-to-end process of managing AI agents throughout their operational life, from planning through testing, deployment, monitoring, governance and decommissioning. AWS Prescriptive Guidance says operationalizing agentic AI means embedding governance, scalability and business alignment, including lifecycle management.
The NIST AI Risk Management Framework 1.0, released 26 January 2023, is voluntary, with four functions: govern, map, measure, manage. NIST says it is being revised. NIST also says the Core's actions are not a checklist, so the stage gates below are this article's own.
How it differs from MLOps, model lifecycle, AI governance and ordinary application release management
IBM separates model management, which focuses on the model, from agent lifecycle management, which covers the full agent system, including prompts, tools, memory, integrations, access control, evaluations and decommissioning. IBM says the practice builds on MLOps.
In this article's framing, model lifecycle work centres on the model; release management, on a build; agent lifecycle management adds a record of authority for one agent. AI agent governance sets organisation-wide policy; the AI governance framework covers its controls.
Why agents need action, tool, identity and autonomy controls in addition to model evaluation
IBM says agents can call tools and automate actions, might produce different outputs for similar inputs and, without proper access control, can become over-permissioned non-human identities.
The NIST glossary defines least privilege, in SP 800-53 Rev. 5 wording, as granting each entity the minimum system resources and authorisations it needs to perform its function. A model evaluation does not show whether a refund tool has a spending limit, so permissions, policy and approval points need their own controls.
The eight stages of an AI agent lifecycle
Each stage ends at a gate: a named owner confirms that specific evidence exists before the agent moves on. IBM organises a practical model around five phases; this article uses eight so that risk classification, approval and change control each get a gate.
1. Propose and register
Record the use case, owner, intended outcome, dependencies and users. This gate is registration, not a full inventory; NIST Govern 1.6 describes mechanisms to inventory AI systems.
2. Classify risk and authority
Rate autonomy, data sensitivity, tool reach, consequence of error and reversibility, then assign a risk level and an authority level such as read-only, draft-only or act-with-approval. To assess AI system risk with a scored method, use a dedicated guide.
These levels are this article's own, separate from OpenBox Trust Tiers, where a lower tier number means more autonomy.
3. Design
Write down the permissions the agent will hold, the policy that limits them, where a human approves and what happens when a tool fails. Architecture choices are covered in AI agent design patterns.
4. Build and evaluate
Test task success, then the controls. IBM's evaluation types include regression, prompt-injection, data-leakage, policy-compliance and human-in-the-loop approval testing. NIST uses TEVV (test, evaluation, verification, validation); Measure 2.1 describes documenting test sets, metrics and tools.
5. Approve and release
The approver reviews the evidence package, states the residual risk and signs off one version with a rollout and rollback plan. NIST Manage 1.4 describes documenting negative residual risks, which it defines as the sum of all unmitigated risks. IBM lists phased rollout, version pinning for models and prompts, and rollback plans among common release practices.
A stage gate is a checkpoint where the owner confirms the evidence exists. A production-readiness review is the record made at gate 5.
6. Operate and monitor
After release, check that the controls work as well as that the agent runs. NIST Measure 2.4 describes monitoring functionality and behaviour in production. IBM lists policy violations, escalation and approval rates and task success rates among signals to watch. The agent incident response postmortem covers detection and containment.
7. Change and reapprove
IBM says prompts, tools, models and policies should be versioned, reviewed and documented, because changes to them can affect behaviour. Treat a change that alters what the agent can do as a new release.
8. Suspend or retire
Suspend, then retire. IBM says decommissioning should include revoking credentials, removing service accounts, preserving logs, archiving evidence and notifying users. NIST Govern 1.7 describes decommissioning AI systems safely. Here, decommissioning means controlled removal from service.
Stage gates, owners and evidence
Owners are roles; the gates are this article's proposal, so adapt both. The 90-day CISO playbook covers programme rollout.
Stage | Accountable owner | Evidence produced | Exit criterion |
|---|---|---|---|
1. Propose and register | Product owner | Registration record: purpose, owner, users, dependencies | Owner named; agent in the system of record |
2. Classify risk and authority | Risk owner | Classification with rationale | Risk and authority levels assigned |
3. Design | Platform lead | Permissions, policy, approval points, failure modes | Security reviews least privilege |
4. Build and evaluate | Platform lead | Test report, including failures | Planned tests run; failures fixed or accepted |
5. Approve and release | Risk owner | Approval record: residual risk, version, rollout, rollback | Named approver signs one version |
6. Operate and monitor | Operations lead | Monitoring reviews: incidents, exceptions, drift | Reviews done on cadence and after events |
7. Change and reapprove | Risk owner | Change record and regression results | Change classified; approval matches class |
8. Suspend or retire | Product owner | Retirement record | Known access paths closed and tested; owners informed |
RACI for product owner, platform, security, privacy, legal, risk and operations
R is responsible, A accountable, C consulted, I informed. Each activity has one A.
Activity | Product owner | Platform | Security | Privacy | Legal | Risk | Operations |
|---|---|---|---|---|---|---|---|
Propose and register | A/R | C | I | I | I | C | I |
Classify risk and authority | R | C | C | C | C | A | I |
Design permissions and policy | C | A/R | R | C | I | I | C |
Build and evaluate | C | A/R | R | R | I | I | C |
Approve and release | R | C | C | C | C | A | R |
Operate and monitor | C | C | C | I | I | I | A/R |
Approve a material change | R | R | C | C | C | A | C |
Suspend or retire | A | R | R | C | C | I | R |
Minimum evidence at each gate: intended controls versus tested controls
An intended control is a design decision, such as a spending limit in the policy. A tested control has a recorded test showing the limit holds. The evidence register below marks each record; approve on tested controls and list controls that are intended but untested as residual risk.
Record | Type | Stage | What it shows |
|---|---|---|---|
Registration record | Intended | 1 | Owner and purpose |
Classification record | Intended | 2 | Risk level, authority level, rationale |
Permission and policy specification | Intended | 3 | What the agent may do, under which policy version |
Test report | Tested | 4 | Which controls held against which cases |
Approval record | Decision | 5 | Residual risk accepted; version approved |
Monitoring review | Observed | 6 | Behaviour and control effectiveness in production |
Change record | Tested | 7 | What changed, its class, the regression result |
Retirement record | Observed | 8 | Access removed; evidence archived |
Production-readiness checklist and approval record
What has to be true before production? Each item below, with evidence.
[ ] Agent registered with a named accountable owner
[ ] Outcome, users and dependencies recorded
[ ] Risk level and authority level assigned, with rationale
[ ] Permissions follow least privilege and are listed by tool, with the policy version
[ ] Human approval points, approvers and tool failure behaviour defined
[ ] Task, safety, security, privacy and adversarial tests run
[ ] Failed cases fixed or accepted in writing
[ ] Residual risk stated and accepted by the named approver
[ ] Version, rollout plan and rollback plan recorded
[ ] Monitoring signals, alert owners and review cadence set
[ ] Suspend procedure and incident contact confirmed
The approval record names the agent, version, policy version, residual risk, approver and review date.
What counts as a material agent change?
A material change alters what an agent can do, what data it touches, who it affects or how far it acts without a human. It invalidates the previous approval until the matching review is done. This is an operating model, not a legal or NIST requirement.
Changes that trigger full review
These changes trigger full reapproval:
A new write permission: create, update, delete, send or spend
A move to a higher-impact use
Access to sensitive data
A new external recipient
Higher autonomy, meaning more decisions made without a human approval point
Tool changes deserve care because agents can combine legitimate tools in unintended ways; new agent-to-agent handoffs (governing handoffs) belong in the same class.
Risk-based review for model, prompt, retrieval, tool, orchestration and policy changes
Other changes get a review sized to the stage 2 risk level:
Change type | Re-check before release |
|---|---|
Model version | Full regression suite; compare results |
Prompt or instruction | Regression suite; adversarial cases on the prompt |
Retrieval source or data | Data-access and grounding tests; privacy review if the data class changes |
Tool or API schema | Permission review; tool-call tests |
Orchestration or routing | Regression suite; failure tests |
Policy or guardrail | Known-good and known-bad inputs; record the policy version |
Decision tree: which review does a change need?
Work down the steps and stop at the earliest yes.
Step | Does the change | If yes |
|---|---|---|
1 | add a write, spend, delete or send capability, or a new external recipient? | Full reapproval at stage 5 |
2 | move to a higher-impact use, or add sensitive data? | Reclassify at stage 2, then reapprove |
3 | remove or loosen a human approval? | Reclassify at stage 2, then reapprove |
4 | alter the model, prompt, retrieval, tool schema, orchestration or policy? | Risk-based review, canary, rollback ready |
5 | leave what the agent can do untouched, such as a copy fix? | Log it; no reapproval |
Regression suite, canary, rollback and change log
A regression suite is the tests re-run after each change. A canary release exposes it to a small traffic share before full rollout. Rollback restores the prior version. The change log records each change, its class, results and approver.
Monitor operating effectiveness after release
Monitoring answers whether the controls work as well as whether the agent works. It shows what happened; permission comes from the approved policy, not from monitoring. See monitoring agents in production.
Task success and quality versus unauthorised actions, interventions, policy misses and drift
Track task success beside control signals: unauthorised actions, human interventions, approval outcomes and goal drift. Sample allowed actions as well as blocked ones: a policy miss should have been stopped and was not, and a sample can reveal it.
Review cadence and event-based reassessment
Set a scheduled review and event triggers at approval. NIST Govern 1.5 describes planning periodic review, including determining its frequency. OpenBox's Compliance & Audit guidance suggests reviewing governance policies and Behavioral Rules quarterly. Triggers include a material change, incident, drift alert or upstream model change.
Worked example: customer-support agent gains refund authority
Hypothetical.
Prior authorisation. The agent answers questions, drafts replies and reads order history. It is read-only, with no payment access.
New classification. Refunds add a spend capability, which triggers full review. The team raises the risk level and sets the authority level to act-with-approval above a limit that finance sets.
Tests. Cover refunds inside and above the limit, the wrong customer, a duplicate refund, a prompt injection asking for a refund and a failing payment tool. Privacy tests check that no other customer's data is revealed.
Approval. The risk owner reads the test report, records the residual risk (for example, eligibility judgement on unusual cases) and approves one version, with human approval gates above the limit.
Staged release. A canary exposes refunds to a small share of conversations at a low limit.
Monitoring and fallback. Reviewers sample allowed refunds and track reversals and interventions. If refunds misbehave, the team pauses the capability or the agent and reverts to draft-only.
Retirement and exit checklist
Retirement is done when every known access path is shown closed and the evidence is kept or deleted according to policy:
[ ] Stop scheduled runs and triggers
[ ] Pause the agent and confirm it starts no new sessions
[ ] Revoke the agent's API keys and signing keys
[ ] Remove tool grants, service accounts and integrations
[ ] Call the agent with the old credentials and record the rejection
[ ] Archive the evidence: registration, approvals, tests, changes, monitoring reviews
[ ] Apply retention and deletion rules after legal and privacy review
[ ] Notify owners and users, and update the catalogue
The rejected call shows the revoked credentials no longer work on that path. It does not show every path is closed; confirm tool grants and other integrations separately.
How OpenBox fits the lifecycle
OpenBox, an AI agent governance platform, organises governance as a Trust Lifecycle of five phases: Assess, Authorize, Monitor, Verify and Adapt. It complements build, testing, legal and change-management systems. Its Agent Lineage documentation says lineage is not a build system or artefact registry and does not replace GitHub or your deployment system.
Lifecycle stage | Documented OpenBox capability |
|---|---|
1. Propose and register | Registering Agents creates the agent entity, generates an API key and sets the initial risk profile. Registration also provides a DID and signing key (Agent Identity). |
2. Classify risk and authority | Assess sets a Risk Profile, with documented re-assessment triggers such as changed capabilities. |
3. Design | Authorize covers Guardrails, Policies in OPA Rego and Behavioral Rules, producing one of the governance decisions: ALLOW, CONSTRAIN, REQUIRE_APPROVAL, BLOCK or HALT. |
4. Build and evaluate | A Test Guardrail panel, Rego policy testing and Behavioral Rule test examples; the docs advise testing policies and Behavioral Rules before rollout. Keep your own tooling for task, safety and adversarial evaluation. |
5. Approve and release | Agent Lineage governance snapshots record policy, guardrail and Behavioral Rule version hashes. Release approval stays in your change process. |
6. Operate and monitor | Monitor metrics, Verify drift events, Session Replay and the Approvals queue. Goal alignment needs goal context from your workflow. |
7. Change and reapprove | Compliance & Audit lists Policy Change, Guardrail Change and Risk Configuration Change events; Adapt offers policy suggestions. |
8. Suspend or retire | The Agent Settings Danger Zone offers Pause Agent, and Revoke Agent Access, which is permanent. |
The docs say Revoke Agent Access invalidates all API keys, disconnects active integrations and preserves the agent's data and history for audit, and that the agent cannot be reactivated. Credentials and grants in your own systems fall outside what those docs describe, so revoke them there too.
On evidence, the Attestation documentation describes session events as hashed into a Merkle tree, with the Merkle root digitally signed. OpenBox's documentation describes the audit trail as immutable and the proof certificate as tamper-proof. Read that as detecting changes after signing: a signed record does not show the recorded facts are true or that no events are missing.
To review lifecycle controls, bring an agent release or material change to an OpenBox demo and see how authorisation and evidence fit into your lifecycle. The checklists above are free to copy.
Frequently asked questions
How is agent lifecycle management different from MLOps?
IBM separates model management, which focuses on the model, from agent lifecycle management, which covers the full agent system: prompts, tools, memory, data sources, integrations, access control, audit trails, evaluations, incident response and decommissioning. IBM says it builds on practices including MLOps.
Who owns an AI agent?
Each agent should have one accountable owner, for example the product owner, recorded at registration. Other roles are responsible or consulted at specific stages. This article's RACI is a suggestion, not a standard. NIST Govern 2.1 describes documenting roles and responsibilities.
When should an AI agent be reapproved?
Reapprove when a change alters what the agent can do: a new write permission, a higher-impact use, sensitive data, a new external recipient or higher autonomy. Model, prompt, retrieval, tool, orchestration and policy changes get a risk-based review. This is an operating model, not a legal or NIST requirement.
What is a material change to an AI agent?
A material change alters what an agent can do, what data it touches, who it affects or how far it acts without a human. It invalidates the previous approval until the matching review is done. IBM says changes to prompts, tools, models or policies should be versioned, reviewed and documented.
What evidence should an agent release contain?
A release should contain the registration record, the classification, the permission and policy specification with its version, a test report that includes failed cases, the approval record with residual risk, the rollout and rollback plan, and the monitoring plan. Approve on tested controls.
How do you retire an AI agent?
Pause the agent, then stop scheduled runs, revoke API and signing credentials, remove tool grants and service accounts, and test that the old credentials are rejected. Archive the evidence and apply retention and deletion rules after legal review. IBM also lists notifying users and updating catalogues.
Sources |
OpenBox (PyPI package documentation), "openbox-temporal-sdk-python 1.4.0," https://pypi.org/project/openbox-temporal-sdk-python/1.4.0/, accessed 21 September 2026. NIST, "AI Risk Management Framework," https://www.nist.gov/itl/ai-risk-management-framework, accessed 21 September 2026. NIST AI Resource Center, "AI RMF Core," https://airc.nist.gov/airmf-resources/airmf/5-sec-core/, accessed 21 September 2026. NIST Computer Security Resource Center, "least privilege (Glossary)," https://csrc.nist.gov/glossary/term/least_privilege, accessed 21 September 2026. IBM, "What is agent lifecycle management?," https://www.ibm.com/think/topics/agent-lifecycle-management, accessed 21 September 2026. AWS Prescriptive Guidance, "Operationalizing agentic AI on AWS," https://docs.aws.amazon.com/prescriptive-guidance/latest/strategy-operationalizing-agentic-ai/introduction.html, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Trust Lifecycle," https://docs.openbox.ai/trust-lifecycle, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Assess," https://docs.openbox.ai/trust-lifecycle/assess, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Authorize," https://docs.openbox.ai/trust-lifecycle/authorize, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Guardrails," https://docs.openbox.ai/trust-lifecycle/authorize/guardrails, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Policies," https://docs.openbox.ai/trust-lifecycle/authorize/policies, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Behavioral Rules," https://docs.openbox.ai/trust-lifecycle/authorize/behaviors, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Monitor," https://docs.openbox.ai/trust-lifecycle/monitor, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Verify," https://docs.openbox.ai/trust-lifecycle/verify, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Session Replay," https://docs.openbox.ai/trust-lifecycle/session-replay, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Adapt," https://docs.openbox.ai/trust-lifecycle/adapt, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Governance Decisions," https://docs.openbox.ai/core-concepts/governance-decisions, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Trust Tiers," https://docs.openbox.ai/core-concepts/trust-tiers, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Registering Agents," https://docs.openbox.ai/dashboard/agents/registering-agents, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Agent Settings," https://docs.openbox.ai/dashboard/agents/agent-settings, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Agent Identity," https://docs.openbox.ai/core-concepts/agent-identity, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Agent Lineage," https://docs.openbox.ai/core-concepts/agent-lineage, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Compliance & Audit," https://docs.openbox.ai/administration/compliance-and-audit, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Attestation & Cryptographic Proof," https://docs.openbox.ai/administration/attestation-and-cryptographic-proof, accessed 21 September 2026. OpenBox (docs.openbox.ai), "Approvals," https://docs.openbox.ai/approvals/, accessed 21 September 2026. |

