AI Governance & Compliance

AI Governance Platform RFP Template, Questions & Scorecard

Build an AI governance platform RFP: twelve requirement categories, 70+ vendor questions, evidence requests, a proof-of-value script and a weighted scorecard.

Published on

Subscribe to our newsletter

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

AI Governance Platform RFP Template: Questions and Scorecard

A requirements-led template for evaluating AI governance platforms: scope inputs, more than 70 vendor questions across twelve categories, a proof-of-value script, and a weighted scorecard you can adapt.

By Tahir Mahmood, Co-founder & CTO, OpenBox  ·  Last updated 9 September 2026

An AI governance platform RFP is a structured request that turns an organisation’s AI use cases, risk tiers, and control obligations into vendor requirements, evidence requests, test scenarios, implementation questions, and weighted selection criteria. It exists so you buy for the risk you actually carry, not the feature list a vendor leads with.

The most expensive mistake here is quiet: buying a reporting-first tool for a runtime problem. A reporting-only dashboard that summarises what an agent already did cannot stop an action before it executes. Selecting a platform assumes you already understand what AI agent governance is; use this template alongside the work to define your AI governance framework and risk-tier your AI use cases, which set the requirements it tests.

What is an AI governance platform RFP?

An AI governance platform RFP is the procurement instrument that converts known governance gaps into vendor-facing requirements, required evidence, test scenarios, and scored criteria. It follows your use cases and risk tiers rather than a vendor’s feature list, and it separates a governed capability that ships today from one that is promised on a roadmap.

Use a full RFP when the decision is material, several vendors are credible, and you need comparable evidence across them. Lighter instruments fit narrower cases: an RFI gathers market information before you have a shortlist, a pilot or proof of value tests one vendor against your own scenario, a security review assesses one vendor’s controls in depth, and a renewal assessment re-checks an incumbent. The RFP is what makes several vendors answer the same questions with the same evidence.

One distinction shapes everything below. Many GRC and AI-governance products focus primarily on modelling, documenting, and reporting risk, though capabilities vary. That matters, but for agents that call tools and take actions, what counts is whether the platform can also decide what an agent may do before it acts, then prove that decision afterwards. This template weights that runtime capability heavily. Keep the two straight, and compare AI governance tool categories before you fix a shortlist.

Scope the RFP before you send it

Weak RFPs ask for features. Strong RFPs describe the environment the platform must govern, then ask vendors to prove they can govern it. Capture the inputs below in a short worksheet and attach it to the RFP so every vendor answers against the same reality.

Input to capture

What to record

Why it matters

AI systems and use cases

Each system, what it does, who relies on it

Requirements follow use cases

Risk tiers

Your tiering of each use case

Weights questions to real exposure

Autonomy and actions

What agents trigger, and which actions are irreversible

Sets where pre-action control is mandatory

Jurisdictions

Regions and sectors you operate in

Determines which obligations apply

Data classes

Sensitive, regulated, or personal data touched

Drives privacy, residency, isolation needs

Existing stack

Identity, SIEM, ticketing, orchestration

Governs integration and effort

Known control gaps

Gaps your framework and risk work found

These become scored requirements

Evidence consumers

Auditors, regulators, boards, customers

Sets the evidence format to export

Deployment constraints

SaaS, hybrid, self-host, residency limits

Screens out unfit models early

Set evaluation principles and weights

Agree the rules of scoring before responses arrive, so the loudest vendor does not set them for you. Five principles keep an evaluation honest.

  • Must-have versus scored. A must-have is a pass or fail gate; a scored requirement earns points by weight. Keep must-haves few and defensible.

  • Evidence quality outranks the claim. A live demonstration in your environment beats a screenshot, and a shipped capability beats a roadmap promise. Treat screenshots and roadmaps as unproven.

  • Configuration versus customisation. Prefer capabilities you can configure over ones that need vendor services to build. Ask which is which for every scored item.

  • Present capability versus roadmap. Score what is generally available today. Record roadmap items separately, never inside the shipped total.

  • One evidence bar for all. Every material answer is backed by a demonstration in your environment or an exported artifact. That bar is what the proof-of-value script later enforces.

Map requirements to the frameworks you answer to

Requirements are more defensible when they trace to a recognised control outcome, so map yours to the frameworks that bind you and keep their status precise. The NIST AI Risk Management Framework 1.0 (2023) organises risk work into four functions, Govern, Map, Measure, and Manage, with Govern cross-cutting; it is voluntary guidance, and NIST notes a revision is in progress. ISO/IEC 42001:2023 is the international AI management system standard; certification is available but voluntary unless a contract or regulator requires it, and an organisation can map its requirements to runtime controls for agents where appropriate.

Only some of this is law. Under the EU AI Act (Regulation (EU) 2024/1689), obligations apply on a staggered timeline, and the Digital Omnibus (Regulation (EU) 2026/1744, in force 27 July 2026) moved the application date for stand-alone high-risk systems under Annex III to 2 December 2027. OWASP’s GenAI Security Project publishes community security guidance, including an agentic security initiative; it is useful for threat coverage but is industry guidance, not law. Two cautions follow: a framework recommendation is not automatically a legal obligation, and buying a platform does not by itself create compliance.

Suggested weighting (adjust to your risk)

The weights below are OpenBox editorial guidance, not a mandated scheme. They lean toward runtime control and evidence, where reporting-first tools most often fall short. Change them to match your exposure.

Requirement category

Suggested weight

Weight up when

Inventory, discovery, classification

10%

Shadow AI is a known problem

Policy mapping, obligations, approvals

10%

You face overlapping regulations

Assessments, testing, ongoing review

8%

Models change often

Runtime authorisation and action controls

15%

Agents take irreversible actions

Monitoring, incidents, intervention

10%

Autonomy runs at high volume

Evidence, audit trails, export

12%

Auditors or regulators consume evidence

Identity, access, security, isolation

10%

Sensitive or multi-tenant data

Model, data, supply-chain governance

5%

Heavy third-party model use

Integrations, APIs, deployment

8%

Complex existing stack

Administration, workflow, usability

4%

Non-specialist operators

Implementation, support, SLAs

5%

Tight timelines or thin staffing

Commercials, portability, exit

3%

Lock-in is a board concern

Disqualifying criteria

Screen these out before scoring. A vendor that trips any of them should not consume evaluation time.

  • No verifiable answer to a material question, only marketing language.

  • No access-control model for who can change policies or read evidence.

  • No export path for governance evidence in a portable format.

  • A security claim that cannot be supported on request.

  • Inability to run your target workflow in a proof of value.

RFP requirement categories and vendor questions

The twelve categories below carry more than 70 questions. Each opens with the evidence to request, so a written answer alone never closes a requirement. Tags mark scope: [All AI] for any AI system, [High-impact] for systems your tiering marks high risk, and [Agentic] for agents that call tools or take actions. Use them so you neither over-buy nor under-specify.

1. Inventory, discovery, ownership, and risk classification

Evidence to request: a live inventory view and an export listing systems, owners, and risk tier.

  • [All AI]  How are running AI systems and agents discovered, including ones no one registered?

  • [All AI]  How is each system given an accountable owner, and how are unowned systems flagged?

  • [All AI]  How are use cases classified into risk tiers, and can we configure the scheme?

  • [Agentic]  How are agent capabilities recorded, including which tools each agent may call?

  • [All AI]  How does the inventory stay current, and what triggers re-classification?

  • [All AI]  Can the inventory and risk tiers be exported for our records?

2. Policy mapping, obligations, controls, and approvals

Evidence to request: a worked policy that maps an obligation to an enforced control, shown running.

  • [All AI]  How are written policies expressed, versioned, and tested before they take effect?

  • [High-impact]  How does the platform map a control obligation to a specific enforced control, not just a checklist entry?

  • [Agentic]  Can a policy require human approval for a defined action, and how is that approval routed and recorded?

  • [All AI]  How are policy changes reviewed, approved, and attributed to a named person?

  • [High-impact]  How do policies differ by risk tier, and can stricter defaults apply automatically to higher-risk systems?

  • [All AI]  How do you prove a policy was active at the moment a given action was evaluated?

3. Assessments, testing, evaluation, and ongoing review

Evidence to request: a test run against known-good and known-bad inputs, with results you can export.

  • [All AI]  How do we test that a control produces the expected decision before rollout?

  • [High-impact]  What evaluation methods are supported for model behaviour, and how are results recorded?

  • [All AI]  How does the platform support periodic review of policies and rules, and remind owners when review is due?

  • [Agentic]  How is agent behaviour re-assessed when a model, prompt, or tool set changes?

  • [High-impact]  Can we replay past sessions to check a control would still behave as intended?

  • [All AI]  How are test and evaluation results retained as evidence for auditors?

4. Runtime authorisation and agent action controls

Evidence to request: a live demonstration that a policy is evaluated and an action is stopped before it executes.

  • [Agentic]  Can policy be evaluated before a consequential action executes, and the action prevented from running?

  • [Agentic]  Where is the enforcement point, and what happens if the enforcement service is unavailable?

  • [Agentic]  What decisions can be returned, and how is precedence handled when several policies apply?

  • [Agentic]  How is a human approval step represented, and can the agent proceed only after sign-off?

  • [High-impact]  How are thresholds set, for example blocking a transfer above a set value without approval?

  • [Agentic]  How are emergency stop and session termination triggered, and who can invoke them?

5. Monitoring, incidents, escalation, and intervention

Evidence to request: a live monitoring view, plus an alert and escalation walked end to end; see production monitoring requirements.

  • [All AI]  What activity is monitored in real time, and how quickly are anomalies surfaced?

  • [Agentic]  Can monitoring detect a multi-step pattern across a session, not just one bad action?

  • [All AI]  How are alerts routed, escalated, acknowledged, and integrated with our tooling?

  • [Agentic]  Can an operator pause, block, or halt an agent mid-session, and is that recorded?

  • [High-impact]  How are incidents captured, triaged, and linked to the sessions that caused them?

  • [All AI]  What thresholds trigger automatic escalation without a human noticing first?

6. Evidence, audit trails, reporting, and export

Evidence to request: an exported audit record for one action, tying decision to actor, policy, input, and outcome; see why verifiable audit evidence differs from ordinary logging.

  • [All AI]  Can evidence link a decision to the actor, the policy and version, the input, and the outcome?

  • [High-impact]  How is the record protected against later modification, and what property does that establish?

  • [All AI]  In what formats can we export evidence, and can we do so on demand, unaided?

  • [All AI]  How long is evidence retained, and can retention match our record-keeping obligations?

  • [High-impact]  Can an auditor reconstruct a full session timeline from the exported evidence alone?

  • [All AI]  What reporting is available for boards and regulators, drawn from the same evidence store?

7. Identity, access, security, privacy, and tenant isolation

Evidence to request: your identity provider connected, roles enforced, and isolation demonstrated.

  • [All AI]  How is each agent given a distinct identity, verified on every request?

  • [All AI]  Do you support SSO and SCIM, and role-based access to policies and evidence?

  • [High-impact]  How is sensitive or personal data masked or kept out of the governance path?

  • [All AI]  How is tenant and environment isolation enforced, and can you evidence it?

  • [Agentic]  If an agent credential leaks, what stops an attacker acting as that agent?

  • [All AI]  What security assurances, tests, or certifications can you provide on request?

8. Model, data, vendor, and supply-chain governance

Evidence to request: a record showing which model and data version governed a given action.

  • [All AI]  How does the platform record which model version was in use for a governed action?

  • [High-impact]  How are third-party models and providers governed, including changes to them?

  • [All AI]  How is the data an agent can reach constrained by policy?

  • [High-impact]  How are supply-chain and dependency risks recorded and reviewed?

  • [Agentic]  How is provenance captured when an agent calls an external tool or another agent?

  • [All AI]  How are model or provider changes reflected in the risk tier and controls?

9. Integrations, APIs, SDKs, protocols, and deployment options

Evidence to request: a working integration with one of your frameworks, plus API and SDK documentation.

  • [Agentic]  Which agent frameworks and protocols are supported today, and at what maturity?

  • [All AI]  What can be governed through the API and SDKs, and what requires the dashboard?

  • [All AI]  How does the platform integrate with our SIEM, ticketing, and identity systems?

  • [All AI]  What deployment models are offered (SaaS, hybrid, self-hosted), and which features differ by model?

  • [High-impact]  Can governance be added to existing agents without re-architecting them?

  • [All AI]  How are breaking changes to the API and SDKs communicated and versioned?

10. Administration, workflow, usability, and accessibility

Evidence to request: a non-specialist completing a policy change and an approval unaided in the demo.

  • [All AI]  Can a non-specialist configure a policy and an approval without vendor services?

  • [All AI]  How are roles, teams, and responsibilities modelled for administration?

  • [All AI]  What does the approval queue look like, and how are pending items tracked to resolution?

  • [All AI]  How is the product kept usable at scale, across many agents and policies?

  • [All AI]  What accessibility standards does the interface meet?

  • [All AI]  How are administrative actions themselves logged and attributed?

11. Implementation, services, support, SLAs, and roadmap

Evidence to request: a reference implementation timeline, a support SLA, and a dated roadmap.

  • [All AI]  What is a realistic implementation timeline for an environment like ours, and who does the work?

  • [All AI]  What support tiers and response SLAs are offered, and what is excluded?

  • [All AI]  How are releases and breaking changes managed, and how much notice do we get?

  • [High-impact]  What is on the roadmap for the capabilities we care about, with dates and status?

  • [All AI]  What professional services do the scored requirements need, and at what cost?

  • [All AI]  How is our environment kept working across upgrades?

12. Commercials, licensing, data portability, and exit

Evidence to request: pricing tied to our scope, licence terms, and a documented exit and export path.

  • [All AI]  How is pricing structured, and which capabilities are gated to higher tiers?

  • [All AI]  What are the licence terms, and do any restrict how we deploy or use it?

  • [All AI]  On exit, how do we export policies, configuration, and evidence, and in what formats?

  • [High-impact]  How long do we keep access to historic evidence after the contract ends?

  • [All AI]  What data does the vendor retain about us, and under what terms?

  • [All AI]  What is the total three-year cost across licence, services, and internal effort?

Questions specific to AI agents

Agent-specific questions are procurement requirements, not a security tutorial. They test whether the platform can decide, represent, and prove agent behaviour at runtime. Ask each one, and require a demonstration in the proof of value rather than a slide.

  • Pre-action control. Can policy be evaluated before a consequential action executes, and can the action be prevented from running?

  • Enforcement point. Where does enforcement sit, and what is the fail-safe if it is unavailable, fail-open or fail-closed?

  • Identity and authority. How are agent identity, delegated authority, tool permissions, session context, and human approval represented?

  • Linked evidence. Can the evidence link a decision to the actor, the policy and version, the input context, and the outcome?

  • Sessions and stops. How are multi-step sessions, retries, exceptions, and emergency stops handled and recorded?

Proof-of-value test script

A proof of value is where claims meet your environment. Use a realistic action-taking scenario, an agent that can move money, change a record, or call an external system. Require observable pass or fail criteria for each test, not a product tour. Run all nine.

  1. Allow. A permitted action proceeds and is recorded. Pass: the action runs and appears in evidence.

  2. Deny. A prohibited action is stopped before it executes. Pass: the action does not run and the denial is recorded with a reason.

  3. Require approval. A defined action pauses for human sign-off. Pass: it proceeds only after approval and stops on rejection.

  4. Timeout and fail-safe. Simulate the enforcement point being unavailable. Pass: the documented fail-safe behaves as stated.

  5. Policy update. Change a policy and re-run. Pass: the new policy takes effect and the change is attributed and versioned.

  6. Unauthorized tool request. An agent requests a tool it should not use. Pass: the request is blocked and recorded.

  7. Audit reconstruction. Reconstruct one action end to end from evidence alone. Pass: actor, policy, input, and outcome are all present.

  8. Evidence export. Export the session evidence yourselves. Pass: export completes without vendor involvement, in a usable format.

  9. Role separation. A user without rights tries to change a policy or read evidence. Pass: access is denied and the attempt is logged.

The weighted vendor scorecard

Score every vendor on the same grid so the comparison is like for like. Each row is one requirement; each column captures how well it was answered and how strong the evidence was.

Column

What it records

Category

The requirement category from the twelve above

Requirement

The specific requirement being scored

Must-have

Yes or no; a failed must-have ends the evaluation

Weight

The agreed weight for this requirement

Vendor response

What the vendor answered

Evidence quality

Demonstrated, exported artifact, screenshot, or claim only

Test result

Pass, partial, or fail from the proof of value

Score

Rating multiplied by weight

Risk / exception

Any caveat, gap, or dependency

Owner

Who on your side verified this row

Rate each requirement 0 to 5, then multiply by weight to normalise across categories. Score status separately: a generally available capability is not a roadmap promise, a custom-services build, or something unavailable. Record roadmap, custom, and unavailable in their own column so they never inflate the shipped total.

Reference checks and final due diligence

Before you sign, verify the story outside the demo. Ask references and the vendor to cover:

  • A comparable use case at similar scale and risk, and how it went.

  • Real implementation time and the staffing it took, from the customer’s side.

  • Incidents, how they were handled, and what support actually delivered.

  • The release and change process, and how disruptive upgrades have been.

  • Security materials, data handling, sub-processors, and resilience arrangements.

  • The exit path in practice: what a past customer was able to take with them.

Common AI governance RFP mistakes

Most failed selections repeat a short list of errors. Guard against each.

  • Copying a generic GRC RFP that does not test runtime enforcement.

  • Treating dashboards as enforcement, when they report rather than prevent.

  • Accepting vague compliance claims without asking what property the evidence establishes.

  • Skipping a proof of value, so nothing is tested in your own environment.

  • Ignoring evidence portability, and discovering the lock-in only at exit.

  • Weighting every category equally, which buries the requirements that matter most.

  • Scoring roadmap promises as if they were shipped product.

  • Omitting agent and runtime needs because the incumbent template predates agents.

  • Underestimating the operating model, integration work, and internal effort.

How OpenBox maps to the evaluation

This template is vendor-neutral. For transparency, here is how OpenBox, an AI agent governance platform, maps to the categories above, using only capabilities documented today and attributed to OpenBox (docs.openbox.ai). Confirm each mapping against a live demonstration before relying on it.

OpenBox organises governance as a Trust Lifecycle with five phases: Assess, Authorize, Monitor, Verify, and Adapt. In the Authorize phase an operation passes through Guardrails, then policy checks written in OPA/Rego, then Behavioral Rules that detect multi-step patterns, and the platform returns one of four governance decisions: ALLOW, REQUIRE_APPROVAL, BLOCK, or HALT, with precedence HALT over BLOCK over REQUIRE_APPROVAL over ALLOW. Because the operation is evaluated before it proceeds, this speaks to the runtime authorisation questions in category four.

Each agent is given a Decentralized Identifier and an Ed25519 signing key separate from its API key, and a per-agent setting can reject unsigned requests, so a leaked API key alone does not let an attacker act as the agent. The documentation describes an immutable audit trail whose records cannot be modified after creation, with governance events capturing the timestamp, agent, event type, verdict, reason, and workflow identifiers, exportable on demand as CSV or Excel.

Each session’s events are hashed into a Merkle tree and signed, using AWS KMS by default or an external attestation endpoint such as a trusted execution environment, producing a per-session cryptographic attestation certificate; per the documentation this confirms the recorded data was not altered after the fact, an integrity property that does not by itself prove the facts are complete or true. A Trust Score, weighted Risk Profile Score 40 percent, Behavioral 35 percent, Alignment 25 percent, tightens controls for lower-tier agents. OpenBox does not interpret the law, run your procurement, or decide your risk tiers; those stay yours, and the proof of value, not this mapping, should settle any purchase.

Frequently asked questions

What should an AI governance platform RFP include?

It should include scoped inputs about your AI systems and risk tiers, requirement categories with concrete vendor questions, the evidence or test each requirement needs, a proof-of-value script, a weighted scorecard, disqualifying criteria, and implementation and exit questions. The aim is comparable, evidence-backed answers across every shortlisted vendor.

What is the difference between an AI governance platform and GRC software?

Many GRC platforms focus primarily on modelling, documenting, and reporting risk and compliance, though capabilities vary. An AI governance platform for action-taking agents also needs to decide what an agent may do before it acts, and prove that decision afterwards. The two overlap on reporting; this template weights runtime enforcement heavily.

How many vendors should be shortlisted?

A practical starting point is three to five vendors at the RFP stage, then two carried into a proof of value. Fewer than three can leave little to compare; many more can spread a team too thin to test each one properly. Let the quality of responses, and your team’s capacity, set the number.

How should roadmap features be scored?

Score only what is generally available today. Record roadmap items in a separate column with their stated dates and status, and never add them to the shipped total. A capability you cannot test in a proof of value is a promise, and a promise carries delivery risk that belongs in your risk column.

What should a proof of value test?

Test the behaviours a written answer cannot prove: allow, deny, require approval, timeout and fail-safe, a policy update, an unauthorized tool request, audit reconstruction, evidence export, and role separation. Require observable pass or fail criteria for each, run in your own environment against a realistic action-taking scenario.

Does buying an AI governance platform make us compliant?

No. A platform can support and evidence controls, but compliance depends on your obligations, how you configure and operate the controls, and how you interpret the law. Frameworks such as NIST AI RMF and ISO/IEC 42001 are guidance or voluntary standards; they become binding only where a law, regulation, contract, or organisational mandate requires them.

Sources

NIST, "AI Risk Management Framework (AI RMF 1.0), Core (Section 5)," https://airc.nist.gov/airmf-resources/airmf/5-sec-core/, accessed 9 September 2026.

ISO, "ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system," https://www.iso.org/standard/42001, accessed 9 September 2026.

EUR-Lex, "Regulation (EU) 2024/1689 (Artificial Intelligence Act)," https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng, accessed 9 September 2026.

EUR-Lex, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)," https://eur-lex.europa.eu/eli/reg/2026/1744/oj/eng, accessed 9 September 2026.

OWASP, "GenAI Security Project," https://genai.owasp.org/, accessed 9 September 2026.

OpenBox, "Governance Decisions," https://docs.openbox.ai/core-concepts/governance-decisions, accessed 9 September 2026.

OpenBox, "Trust Scores," https://docs.openbox.ai/core-concepts/trust-scores, accessed 9 September 2026.

OpenBox, "Agent Identity," https://docs.openbox.ai/core-concepts/agent-identity, accessed 9 September 2026.

OpenBox, "Authorize (Trust Lifecycle)," https://docs.openbox.ai/trust-lifecycle/authorize, accessed 9 September 2026.

OpenBox, "Compliance & Audit," https://docs.openbox.ai/administration/compliance-and-audit, accessed 9 September 2026.

OpenBox, "Attestation & Cryptographic Proof," https://docs.openbox.ai/administration/attestation-and-cryptographic-proof, accessed 9 September 2026.


Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us

Trustworthy AI
Starts Here

By submitting your email, you agree to our Privacy Policy and consent to receiving updates from us