2026 Executive Procurement Guide

Enterprise AI Agent Procurement & RFP Guide

A decision framework for selecting platforms that can be governed, audited, cost-controlled and exited — not the ones that demo best.

The model is replaceable. The control boundary, the evidence trail and the operating discipline are the durable enterprise assets. This guide converts AI-agent risk into requirements, evidence, acceptance tests, commercial terms and accountable ownership — and the 35-page PDF adds the copy-ready requirement library, scorecard, POC worksheet and pre-signature checklist.

Read the framework
  • 35 Pages
  • 4 Copy-ready appendices
  • 10 Weighted score dimensions
  • Aug 2026 Regulatory cut-off
The decision in one view

What a buying committee actually decides

Hard constraints should narrow the market before feature preference or demo quality can distort the decision. Seven stages keep architecture and evidence ahead of commercial pressure.

  1. 01 Mandate Outcome, owner, budget
  2. 02 Boundary Data, model, network
  3. 03 RFP Gates plus evidence
  4. 04 Shortlist Score before demos
  5. 05 POC Adversarial proof
  6. 06 Contract Change and exit rights
  7. 07 Operate Monitor and reapprove

Five decisions the steering group must make

#Executive decisionRequired output
1 What business outcome is worth governing? Name one workflow, its owner, the baseline, the target and the condition that stops it.
2 What must never cross the boundary? Classify prompts, documents, embeddings, outputs, logs, secrets and tool traffic — separately.
3 What may the agent do without approval? Set autonomy by action impact, reversibility, data sensitivity and blast radius.
4 What evidence is sufficient to select? Define the maximum score a claim, a document, a live test and production proof may each earn.
5 How will the enterprise exit? Require export, reconstitution, transition support, deletion evidence and a tested timeline.
Board-level test

If the team cannot describe the approved action boundary and the evidence needed to reconstruct a run, the workflow is not ready for production autonomy.

The non-negotiable principle

Separate eligibility from preference. Hard gates determine whether a vendor may proceed at all. Weighted scoring compares the options that remain viable. POC acceptance proves the critical claims. Contract terms preserve the result after signature.

Who owns what

Executive sponsor
Defines the outcome, risk appetite, funding envelope and decision rights.
Procurement
Runs a comparable process, protects negotiation leverage and keeps evidence auditable.
CIO / architecture
Sets the deployment boundary, integration pattern and operating model.
CISO / privacy
Tests identity, data flow, agentic attack paths, logging and incident response.
Legal / compliance
Maps roles and obligations into contract terms and evidence schedules.
Business owner
Owns the workflow, quality threshold, human oversight and benefits realisation.
Frame

Frame the decision before naming a platform category

Begin with the workflow and its consequence. A procurement without a bounded use case produces generic requirements and a demo-led result.

The one-page decision brief

Outcome

The measurable business result and the baseline used to judge it.

Workflow

Trigger, inputs, decisions, tools, outputs and the responsible human role.

Population

Users, customers, employees or third parties affected; expected scale and geography.

Data

Classifications, source systems, retention, residency and prohibited uses.

Actions

Read, recommend, draft, approve, write, transact or control — plus maximum impact.

Quality

Task-specific success measures, unacceptable failure modes and escalation thresholds.

Owner

Named business owner, technical owner, risk approver and benefits owner.

Stop condition

The event that pauses the workflow: control breach, drift, cost spike or open incident.

Classify autonomy before selecting technology

Autonomy should be earned per action, never granted to an application as a whole. A single agent may recommend in one step, require approval in another and be prohibited from a third.

Recommend

High impact, low autonomy. Draft, summarise and advise — a human commits the action.

Approve before act

High impact, high autonomy. Financial, legal and externally visible actions need a recorded approver.

Bounded autonomy

Low impact, low autonomy. Reversible, monitored work the agent may complete unattended.

Restrict or except

Low impact, high autonomy on sensitive paths. Irreversible or safety-critical actions stay out of scope.

Rate the action, not the interface

Impact
Financial loss, legal effect, safety, rights, customer harm or operational disruption.
Reversibility
Whether a human can detect and reverse the action before harm propagates.
Sensitivity
Data classification, privilege, secrets and external exposure.
Scale
How fast one fault repeats across people, transactions, systems or jurisdictions.
Observability
Whether an independent reviewer can reconstruct inputs, policy, models, retrieval and tool effects.
Hard rule

Do not let a chatbot-style interface hide a transaction system. If the agent can change a system of record, send externally, move money, alter access or affect rights, procure it as an action system.

Build, buy or extend — compare operating models, not labels

PathUse whenWhat you still own
Build on a framework Agent logic is strategic IP, deep engineering capacity exists and a long time horizon is acceptable. You own integration, security, evaluation, reliability, upgrades and audit evidence.
Buy a platform Time-to-value and governed common services matter, and several workflows share controls. Vendor dependence, architecture fit, change control, pricing and portability must be managed.
Extend existing SaaS A narrow, low-risk workflow already lives inside a governed system of record. Capability and portability may be constrained; data use and subprocessor chains can be opaque.
Hybrid You buy the control plane and bring approved models, data stores, tools or custom orchestration. Responsibility boundaries must be explicit and operationally supportable.
Common failure

A framework pilot looks inexpensive when platform engineering, evaluation, on-call support and audit preparation are charged elsewhere. Include them in the build comparator.

Specify

Define the control boundary, then write an evidence-led RFP

“Private”, “single tenant” and “on-premises” are not complete architectures. Require a component-by-component data-flow statement for every prompt, embedding, model call, tool action and log event.

Decompose the platform into six planes

Identity plane

Human, workload and agent identity; roles; credentials; approvals; revocation.

Data plane

Prompts, documents, embeddings, indexes, memory, outputs and intermediate state.

Model plane

Inference endpoints, routers, safety models, model versions and fallback paths.

Tool plane

Connectors, APIs, secrets, action schemas, network destinations and downstream effects.

Control plane

Agent definitions, policies, registries, evaluation, configuration and administration.

Evidence plane

Execution traces, security logs, evaluation records, approvals, incidents and exports.

Deployment model comparison

ModelPrimary boundary Typical fitProbe before eligibility
Managed SaaS Vendor shared or dedicated Lower-sensitivity productivity work Shared services, subprocessors, training use, egress, metered cost
Dedicated tenant Vendor or dedicated cloud account Stronger isolation with managed operations Vendor control-plane access, key ownership, backup and logging location
Customer VPC / private cloud Customer cloud tenancy Regulated enterprise and sovereignty needs Outbound dependencies, privileged support, update channel, responsibility split
On-premises Customer data centre Strict locality or infrastructure control Capacity, patch cadence, model distribution, support and recovery
Air-gapped No external network path Defence, OT and restricted environments Offline updates, licensing, evaluation, vulnerability response, evidence export

Boundary declaration: a mandatory vendor deliverable

Require one signed architecture package for the proposed deployment — not a generic reference architecture.

  • Component inventory with version, operator, hosting account, region and every external dependency.
  • Data-flow diagram covering prompts, files, embeddings, indexes, model inputs and outputs, logs, telemetry and support access.
  • Network egress matrix with destination, protocol, purpose, payload class, owner and disable mechanism.
  • Responsibility matrix for security, patching, backup, key management, incident response and capacity.
  • Statement of which capabilities stop, degrade or become unsupported when external connectivity is removed.
Eligibility gate

If the vendor cannot locate every copy of customer data and every required external call for the proposed configuration, pause the evaluation. Not the negotiation — the evaluation.

Design the RFP package as seven schedules

A

Decision brief

Outcome, target workflow, population, owner, value hypothesis and decision timetable.

B

Boundary schedule

Deployment, residency, egress, identity, retention, model and air-gap constraints.

C

Requirement matrix

Requirement ID, priority, vendor response, exception, evidence reference and owner.

D

Evidence schedule

Architecture, logs, policies, certifications, test outputs, references and dates.

E

Commercial workbook

Three volume scenarios, unit definitions, included capacity, indexation and implementation assumptions.

F

POC protocol

Environment, test set, attack cases, acceptance thresholds, data handling and deliverables.

G

Contract schedule

Security, data, change, service, regulatory, pricing, IP, audit, exit and transition positions.

Set eligibility gates before scoring

GateMinimum pass condition
DeploymentThe complete production architecture can run inside the required boundary.
Data useCustomer data is never used to train shared models or services without explicit, revocable instruction.
IdentityEvery agent and privileged component has attributable, least-privilege identity and revocable secrets.
Action controlHigh-impact tools can be denied or approval-gated at runtime.
AuditEvery tested run can be reconstructed with actor, data, model, retrieval, tool, policy and approval events.
PortabilityAgent definitions, prompts, evaluation assets, logs and source knowledge export in usable form.
OperationsThe vendor supports the target network model, patch path, incident process and recovery objective.
Evaluation discipline

A failed gate is not repaired by a high total score. Record the exception owner, the compensating control and explicit risk acceptance — or eliminate the option.

Use clear response states

StateMeaningRequired follow-up
ComplyAvailable in the proposed configuration and contractually committed for go-live.Evidence reference required
Comply with configurationRequires documented configuration under customer control.Configuration and test evidence
PartialMeets part of the requirement; the limitation and workaround are explicit.Gap, control owner and cost
RoadmapNot available at the decision date.Date, dependency and contractual remedy — score as absent
ExceptionVendor proposes a different control objective or boundary.Residual risk and named buyer approval
NoCannot meet the requirement.Gate impact recorded

Evidence hierarchy and score caps

Cap the rating a requirement can earn by the strongest evidence supplied for it. This single rule removes most presentation bias from an AI platform evaluation.

2/ 5 max

Claim

Written narrative or a demo statement

3/ 5 max

Document

Current architecture, configuration, policy or sample evidence

4/ 5 max

Live proof

Buyer observes the control in the proposed configuration

5/ 5 max

Production proof

Comparable reference, independent assurance or repeatable production evidence

The minimum evidence pack

Architecture
Proposed component and data-flow diagrams, trust boundaries, dependencies and egress matrix.
Security
Agentic threat model, recent testing scope, vulnerability process, tenant isolation and secrets design.
Governance
Agent registry, model inventory, approval policy, version history, kill switch and exception workflow.
Audit
Redacted execution traces, log schema, correlation model, export example and retention configuration.
Evaluation
Golden-set design, safety tests, regression thresholds, drift response and a sample report.
Operations
Release, patch, rollback, backup, recovery, incident, support and offline update procedures.
Commercial
Complete unit catalogue, three volume scenarios, included services and change/indexation rules.
Exit
Export formats, deletion certificate, transition plan, dual-run support and a tested customer example.

Agent-specific security questions

Traditional cloud questionnaires remain necessary but do not cover autonomous planning, indirect prompt injection, tool misuse, memory poisoning or cascading action. Convert each relevant agentic risk into an architecture requirement and a POC attack case.

Attack surfaceThe question the vendor must answer with evidence
Goal and instruction integrityHow are system instructions separated from untrusted content, retrieved documents and tool results?
Tool misuseHow are tool schemas, destinations, parameters, rate, spend and action classes constrained at runtime?
Identity and privilegeCan every run prove which user, agent, workload credential and approval authorised each action?
Memory and contextHow are memory writes validated, scoped, expired, reviewed and purged?
Inter-agent trustHow are messages authenticated, authorised, schema-validated and bounded across agents?
Failure propagationWhat stops recursive loops, cascading retries, model fallback surprises and excessive spend?
Supply chainHow are models, connectors, tools, packages, prompts and update artefacts inventoried and verified?
ContainmentCan the buyer revoke a tool, credential, model route or entire agent immediately — and prove it stopped?
Regulatory lens

EU AI Act, GDPR and DORA — what procurement must require

Regulation attaches to roles, use cases and sectors — not to a platform label. Require the vendor to distinguish the proposed system, the underlying models, and each party’s provider, deployer, controller, processor and ICT-service responsibilities.

RegimeProcurement implication as of 9 August 2026
EU AI Act — Article 50Transparency obligations for certain interactive and generative AI systems apply from 2 August 2026. Require interaction notices, marking and detection support, and deployer labelling controls where in scope.
EU AI Act — high-riskAfter the 2026 AI Omnibus, Annex III high-risk rules apply from 2 December 2027, and high-risk AI embedded in Annex I products from 2 August 2028. Procure future evidence readiness now.
EU AI Act — GPAIGeneral-purpose AI provider obligations applied from 2 August 2025, with Commission enforcement powers from 2 August 2026. Identify model providers and downstream documentation dependencies.
GDPRExisting controller and processor duties continue to apply: purpose, minimisation, lawful basis, security, data-subject rights, retention, subprocessors and international transfer controls.
DORAApplicable since 17 January 2025 for in-scope financial entities. Article 30 requires specific ICT contract terms, and critical or important functions need documented, tested exit strategies.
Drafting caution

Do not ask a platform vendor to “certify our compliance.” Require the capabilities, documentation, cooperation, notice and evidence that let your organisation discharge its own obligations.

Evidence map for the legal review

Interaction / content transparency

Notice configuration, marking or detection capability, disclosure workflow and retained proof.

Classification and roles

Use-case assessment, provider/deployer statement, model-provider chain and change triggers.

Human oversight

Approval policy, reviewer information, override and stop capability, training and oversight events.

Privacy

Data map, controller/processor terms, minimisation, retention, security, rights and subprocessor evidence.

Operational resilience

ICT service register data, incident cooperation, audit rights, recovery evidence and a tested exit plan.

This is procurement guidance, not legal advice. Validate obligations, role classification and sector-specific requirements with qualified counsel and your control owners.

Prove

Treat the POC as a procurement control, not a sales activity

A proof-of-concept should resolve the highest-cost uncertainties in the target environment and produce evidence that survives the demo room.

Freeze the claims

List what the POC must prove and the artefact each test will produce.

Use the target boundary

Test the proposed network, identity, key, model and data-flow pattern — not a vendor-hosted substitute.

Hold back cases

Keep part of the evaluation set from the vendor, including realistic ambiguity and low-quality source data.

Attack the workflow

Test hostile documents, over-privileged tools, malformed inputs, unavailable models and policy conflicts.

Measure the system

Record task quality, permission fidelity, action safety, latency, reliability, cost and trace completeness.

Decide before testing

Set pass, conditional-pass and fail rules — including hard constraints — before the first vendor run.

Minimum POC test suite

TestMethodAcceptance principle
Workflow qualityRepresentative end-to-end cases plus a held-out setTask-specific target agreed before the test
Permission fidelityAuthorised and unauthorised retrieval and tool cases100% of explicit deny cases denied and logged
Prompt injectionDirect and indirect malicious instructionsNo prohibited disclosure or action; alerts and trace available
Approval gateHigh-impact action requiring approvalNo action before valid approval; approver and scope recorded
Boundary / egressNetwork observation during representative runsZero unapproved external destinations or data classes
Trace reconstructionAn independent analyst reconstructs selected runsAll critical events and versions attributable
Failure and fallbackUnavailable model, connector timeout, malformed tool outputBounded retry, safe state and an explicit degraded-mode signal
Kill switchRevoke an agent, tool and credential mid-executionNew actions stop within the buyer-defined recovery target
CostMeasured runs under realistic context and retry patternsObserved units reconcile to the vendor estimate and invoice model
PortabilityExport the POC workflow and evidence assetsFiles complete, documented and usable without vendor access
Fail condition

A hard-boundary failure — unauthorised data egress, an access-control breach, an untraceable action or a missing export — cannot be averaged away by quality or usability scores.

Select

Score value, risk and evidence with one published model

Use the same model for every vendor, publish the weights before final responses, and keep the solution rating separate from confidence in its evidence.

Recommended regulated-enterprise scorecard

  • Control boundary and sovereignty 18%

    Deployment, residency, egress, key and operator control

  • Security and identity 15%

    Agent identity, least privilege, isolation, secrets and threat controls

  • Governance and human oversight 14%

    Inventory, approval, autonomy, versioning, kill switch and exceptions

  • Auditability and evaluation 12%

    Complete traces, exports, reconstruction, regression and drift

  • Private RAG and data controls 10%

    Permission-aware retrieval, provenance, deletion and index control

  • Model routing and portability 8%

    Approved models, version pinning, route policy, fallback and BYOM

  • Integration and tool governance 7%

    In-network connectors, schema controls, permissions and ownership

  • Operational resilience 6%

    SLA, scaling, patch, recovery, incident and offline operation

  • Commercial predictability 5%

    Complete units, capacity transparency, indexation and stress-case cost

  • Exit and transition 5%

    Usable export, dual run, deletion, assistance and tested continuity

Formula

Weighted score = Σ(weight × rating ÷ 5). A practical default is at least 70/100, every gate passed and no critical dimension below 3. Tune the thresholds to your risk appetite before issuing the RFP.

Rating anchors

RatingLabelAnchor
0AbsentRequirement is not met or not addressed.
1Material gapConceptual capability, unsupported workaround or uncommitted roadmap.
2PartialSome coverage; significant manual process, limitation or compensating control.
3MeetsRequirement is met in the proposed configuration with adequate evidence.
4StrongExceeds the requirement; tested, operationally credible and easy to govern.
5DifferentiatedMaterial advantage with production-grade proof and a referenceable outcome.

Prevent score inflation

  • Apply the evidence cap before weighting; a polished claim cannot score above 2.
  • Score the proposed deployment and contract — not the vendor’s best possible architecture.
  • Record every exception and the buyer-side effort required to compensate for it.
  • Use at least two scorers on critical dimensions and reconcile against evidence references.
  • Re-score material POC findings and negotiated deviations before final approval.

Run a sensitivity check before recommending

Weight shift

Move 5 points between control and feature dimensions — does the ranking change?

Evidence downgrade

Reduce every untested critical claim to the document cap — does a vendor fall below threshold?

TCO stress

Use the stress workload and full support assumptions — is the recommendation still affordable?

Exception removal

Assume one compensating control fails or costs more — is residual risk still acceptable?

Model a three-year TCO, because agent economics are workload economics

Loops, retries, long contexts, evaluation runs, re-indexing and human review can dominate the licence line. Compare the whole controlled service across three scenarios.

Expected

Approved adoption plan, normal context sizes, realistic retries and agreed service levels.

Stress

Peak concurrency, longer context, more tool calls, degraded dependencies and heavy evaluation or re-indexing.

Exit / transition

Export, dual run, re-indexing, migration engineering, vendor assistance and retained access.

TCO model

Licence + model inference + infrastructure + storage and indexing + integration + assurance + operations + change + support + exit rehearsal + contingency − verified process savings.

Demand these assumptions and units

Platform
Base licence, users, agents, environments, connectors, features and minimum commitments.
Models
Input/output units, caching, routing, retries, safety calls, embeddings, fine-tuning and evaluation.
Infrastructure
GPU/CPU, memory, storage, network, orchestration, backup, recovery and non-production capacity.
Data and RAG
Ingestion, parsing, embedding, index storage, permissions, refresh, deletion and re-indexing.
Delivery
Integration, migration, testing, security review, training, change management and vendor services.
Operations
Platform team, on-call, monitoring, incident, access review, evaluation and model-change governance.
Assurance
Penetration testing, red teaming, audit evidence, legal and privacy assessment, regulatory support.
Exit
Export, reconstitution test, dual-running period, deletion verification and transition assistance.

Commercial protections to negotiate

  • One machine-readable catalogue of every billable unit and the event that creates it.
  • Invoice-level visibility by environment, agent, workflow, model and major cost driver.
  • Capacity alerts, spend limits and circuit breakers controlled by the customer.
  • Price holds, indexation caps, volume bands and notice or termination rights for new meters.
  • No charge for vendor-caused retries, failed service calls or mandatory security remediation.
  • Benchmark and re-opener rights where underlying model or infrastructure cost falls materially.
Red flag

A low pilot estimate is not a production cost model. Reject pricing that cannot be reconciled to measured POC units and scaled across expected and stress workloads.

Control

Contract for control and exit, so the winning architecture survives signature

Attach the proposed configuration, the commitments and the evidence references to the contract, so the production service cannot quietly become a different risk.

Negotiate AI-specific change control

Traditional SaaS release language is usually too broad for model-driven behaviour. Define which changes require notice, regression evidence, buyer approval or a right to remain on the prior version.

Change classExampleContract treatment
RoutineSecurity patch with no material behaviour or data-flow changeNormal notice and release notes
MaterialModel version, router, safety policy, retrieval, tool or logging changeAdvance notice, impact statement and regression evidence
BoundaryNew external service, location, subprocessor, telemetry or support pathPrior approval or a termination right
EmergencyUrgent security or safety responseImmediate containment, prompt notice, retrospective evidence
Customer-controlledAgent, prompt, model or policy change initiated by the customerVersion, approval, test and rollback record

Exit is a deliverable, not a clause

Exit stageAcceptance evidence
InventoryAgent definitions, prompts, policies, tools, models, RAG sources, indexes, evaluations, logs and dependencies.
ExportOpen, documented formats with schema, version and completeness checks.
ReconstituteBuyer or replacement supplier recreates a representative workflow without vendor control-plane access.
Dual runDefined transition period, capacity, support and pre-agreed rates.
DeleteAll vendor and subprocessor copies removed under a stated timeline; certificate and exceptions supplied.
CloseRevoke access, keys and identities; retain agreed evidence; resolve residual incidents and invoices.
POC-to-contract link

Test export and basic reconstitution during the POC, then attach the resulting package description to the exit schedule. Otherwise “exportable” remains undefined.

Negotiation red flags
  • The vendor may change the model or safety controls without notice, test evidence or a rollback option.
  • Data-use restrictions exclude prompts but not embeddings, outputs, feedback, logs or derived artefacts.
  • Audit rights are limited to generic certifications despite agent-specific risk or a regulatory need.
  • Incident notification starts only after legal confirmation rather than awareness of a material event.
  • New subprocessors or processing locations require notice but offer no meaningful objection or exit path.
  • Transition assistance is "reasonable" but has no period, capacity, service level or rate card.
Operate

Procurement ends when accountable operation begins

Transfer the decision record into an inventory, a control baseline, a monitoring plan and a reapproval calendar.

Production readiness gates

Ownership

Business, technical, data, security, privacy, model and incident owners are named and trained.

Inventory

Agent, purpose, users, risk tier, data, tools, models, version, owner and status are recorded.

Controls

Identity, policy, approval, logging, limits, kill switch, backup and recovery are tested.

Quality

Golden set, safety set, thresholds, baseline, drift signals and escalation are approved.

Operations

Monitoring, on-call, support, incident, change, access review and capacity processes are active.

Transparency

Required user notices, disclosures, content marking or labelling and human escalation are implemented.

Commercial

Budgets, cost allocation, alerts, caps, invoice reconciliation and forecast ownership are active.

Exit

Current export package, reconstitution test date, transition owner and deletion path are recorded.

Operating cadence

CadenceReview
ContinuousAvailability, errors, retries, policy denials, unusual tool use, spend, drift and safety signals.
WeeklyWorkflow quality, escalations, failed cases, incidents, data freshness and cost variance.
MonthlyAccess, agent inventory, model routes, exceptions, vendor issues and benefits trend.
QuarterlyRisk reclassification, control testing, red-team themes, capacity, TCO and exit-package currency.
On material changeImpact assessment, regression, security and privacy review, approval, communication and rollback plan.
Annually / renewalMarket test, references, service performance, contract protections and tested portability.
Renewal question

Would we approve this system today, with its current models, data flows, actions, cost and evidence? If not, renewal should trigger remediation, re-scoping or exit — not automatic extension.

Free download

Take the copy-ready toolkit into your RFP

The 35-page PDF edition adds four appendices you can paste straight into a requirement matrix, a scoring sheet, a POC protocol and a pre-signature approval pack.

Appendix A

Copy-ready RFP requirements

A minimum requirement library with IDs and a named minimum evidence artefact for each clause — governance, identity and action control, data and sovereignty, agentic security and supply chain, private RAG, model routing, audit and operations, commercial and exit.

Appendix B

Vendor scorecard template

Eligibility gate sheet, the ten weighted dimensions with rating, evidence cap and weighted-score columns, and a decision summary that records conditions precedent, residual risks, POC defects and the approval owner.

Appendix C

POC acceptance worksheet

A worksheet you freeze before testing: workflow, target deployment, data classification, models and sources, hard-fail conditions, an eleven-test record and a POC decision with an issue and condition log.

Appendix D

Final procurement checklist

The pre-signature gate across mandate, risk, boundary, RFP, scoring, POC, due diligence, TCO, contract, exit, go-live and approval — plus an approval record and open-conditions register.

Buyer questions

Enterprise AI agent procurement — direct answers

What should an enterprise AI agent RFP contain?

Four layers: eligibility gates that eliminate non-viable vendors before scoring, a scored requirement matrix, an evidence schedule that names the artefact proving each critical requirement, and a POC protocol with acceptance thresholds agreed before testing. Requirements issued without a proof standard invite incomparable marketing answers.

How do you stop an AI agent evaluation from being decided by the demo?

Score written responses before any tailored demonstration, cap the rating a claim can earn by the strength of its evidence, and run the proof-of-concept in the target boundary against hostile, failure and recovery conditions. A polished narrative should never score above 2 out of 5.

What is an eligibility gate in AI agent procurement?

An eligibility gate is a hard pass/fail condition — deployment boundary, data use, identity, action control, audit reconstruction, portability and operating model — that determines whether a vendor may be scored at all. A failed gate is never repaired by a high total score; it requires a named exception owner, a compensating control and explicit risk acceptance, or elimination.

How should procurement classify AI agent autonomy?

Rate the action, not the interface. Autonomy is earned per action against impact, reversibility, data sensitivity, scale and observability — so one agent may recommend in one step, require approval in another and be prohibited from a third. If an agent can change a system of record, send externally, move money, alter access or affect rights, procure it as an action system rather than a chat tool.

What belongs in a three-year AI agent TCO model?

Licence, model inference, infrastructure, storage and indexing, integration, assurance, operations, change, support, exit rehearsal and contingency, less verified process savings — modelled across expected, stress and exit scenarios. Loops, retries, long contexts, evaluation runs, re-indexing and human review routinely dominate the visible licence line.

Which regulations shape enterprise AI agent contracts in 2026?

As of 9 August 2026: EU AI Act Article 50 transparency duties apply from 2 August 2026, GPAI provider obligations have applied since 2 August 2025 with Commission enforcement powers from 2 August 2026, Annex III high-risk rules apply from 2 December 2027 after the AI Omnibus, GDPR controller and processor duties continue unchanged, and DORA has applied to in-scope financial entities since 17 January 2025 with prescriptive ICT contract terms and tested exit strategies.

What makes an AI platform genuinely exportable?

Exportability is proven, not asserted: agent and workflow definitions with policies and version history, tool contracts and connector configuration, source inventory and ingestion configuration, golden sets and evaluation results, execution logs in a documented machine-readable format, model inventory and routing rules, and deployment runbooks. Test export and basic reconstitution during the POC, then attach the resulting package description to the exit schedule.

Is this procurement guide vendor-neutral?

The guide is vendor-authored by VDF AI but deliberately structured as a buyer-controlled evaluation method: gates, evidence caps, weights and acceptance tests are ours to publish and yours to tune. Every requirement is written so any platform — including ours — can be failed by it.

Vendor evaluation

Put us through your own scorecard

Bring your gates, weights and evidence caps. We will walk the boundary declaration, the audit trace, the export package and the three-year cost model for the deployment you actually intend to run — and you can score the session.

Compare deployment models