Outcome
The measurable business result and the baseline used to judge it.
A decision framework for selecting platforms that can be governed, audited, cost-controlled and exited — not the ones that demo best.
The model is replaceable. The control boundary, the evidence trail and the operating discipline are the durable enterprise assets. This guide converts AI-agent risk into requirements, evidence, acceptance tests, commercial terms and accountable ownership — and the 35-page PDF adds the copy-ready requirement library, scorecard, POC worksheet and pre-signature checklist.
Hard constraints should narrow the market before feature preference or demo quality can distort the decision. Seven stages keep architecture and evidence ahead of commercial pressure.
| # | Executive decision | Required output |
|---|---|---|
| 1 | What business outcome is worth governing? | Name one workflow, its owner, the baseline, the target and the condition that stops it. |
| 2 | What must never cross the boundary? | Classify prompts, documents, embeddings, outputs, logs, secrets and tool traffic — separately. |
| 3 | What may the agent do without approval? | Set autonomy by action impact, reversibility, data sensitivity and blast radius. |
| 4 | What evidence is sufficient to select? | Define the maximum score a claim, a document, a live test and production proof may each earn. |
| 5 | How will the enterprise exit? | Require export, reconstitution, transition support, deletion evidence and a tested timeline. |
If the team cannot describe the approved action boundary and the evidence needed to reconstruct a run, the workflow is not ready for production autonomy.
Separate eligibility from preference. Hard gates determine whether a vendor may proceed at all. Weighted scoring compares the options that remain viable. POC acceptance proves the critical claims. Contract terms preserve the result after signature.
Begin with the workflow and its consequence. A procurement without a bounded use case produces generic requirements and a demo-led result.
The measurable business result and the baseline used to judge it.
Trigger, inputs, decisions, tools, outputs and the responsible human role.
Users, customers, employees or third parties affected; expected scale and geography.
Classifications, source systems, retention, residency and prohibited uses.
Read, recommend, draft, approve, write, transact or control — plus maximum impact.
Task-specific success measures, unacceptable failure modes and escalation thresholds.
Named business owner, technical owner, risk approver and benefits owner.
The event that pauses the workflow: control breach, drift, cost spike or open incident.
Autonomy should be earned per action, never granted to an application as a whole. A single agent may recommend in one step, require approval in another and be prohibited from a third.
High impact, low autonomy. Draft, summarise and advise — a human commits the action.
High impact, high autonomy. Financial, legal and externally visible actions need a recorded approver.
Low impact, low autonomy. Reversible, monitored work the agent may complete unattended.
Low impact, high autonomy on sensitive paths. Irreversible or safety-critical actions stay out of scope.
Do not let a chatbot-style interface hide a transaction system. If the agent can change a system of record, send externally, move money, alter access or affect rights, procure it as an action system.
| Path | Use when | What you still own |
|---|---|---|
| Build on a framework | Agent logic is strategic IP, deep engineering capacity exists and a long time horizon is acceptable. | You own integration, security, evaluation, reliability, upgrades and audit evidence. |
| Buy a platform | Time-to-value and governed common services matter, and several workflows share controls. | Vendor dependence, architecture fit, change control, pricing and portability must be managed. |
| Extend existing SaaS | A narrow, low-risk workflow already lives inside a governed system of record. | Capability and portability may be constrained; data use and subprocessor chains can be opaque. |
| Hybrid | You buy the control plane and bring approved models, data stores, tools or custom orchestration. | Responsibility boundaries must be explicit and operationally supportable. |
A framework pilot looks inexpensive when platform engineering, evaluation, on-call support and audit preparation are charged elsewhere. Include them in the build comparator.
“Private”, “single tenant” and “on-premises” are not complete architectures. Require a component-by-component data-flow statement for every prompt, embedding, model call, tool action and log event.
Human, workload and agent identity; roles; credentials; approvals; revocation.
Prompts, documents, embeddings, indexes, memory, outputs and intermediate state.
Inference endpoints, routers, safety models, model versions and fallback paths.
Connectors, APIs, secrets, action schemas, network destinations and downstream effects.
Agent definitions, policies, registries, evaluation, configuration and administration.
Execution traces, security logs, evaluation records, approvals, incidents and exports.
| Model | Primary boundary | Typical fit | Probe before eligibility |
|---|---|---|---|
| Managed SaaS | Vendor shared or dedicated | Lower-sensitivity productivity work | Shared services, subprocessors, training use, egress, metered cost |
| Dedicated tenant | Vendor or dedicated cloud account | Stronger isolation with managed operations | Vendor control-plane access, key ownership, backup and logging location |
| Customer VPC / private cloud | Customer cloud tenancy | Regulated enterprise and sovereignty needs | Outbound dependencies, privileged support, update channel, responsibility split |
| On-premises | Customer data centre | Strict locality or infrastructure control | Capacity, patch cadence, model distribution, support and recovery |
| Air-gapped | No external network path | Defence, OT and restricted environments | Offline updates, licensing, evaluation, vulnerability response, evidence export |
Require one signed architecture package for the proposed deployment — not a generic reference architecture.
If the vendor cannot locate every copy of customer data and every required external call for the proposed configuration, pause the evaluation. Not the negotiation — the evaluation.
Outcome, target workflow, population, owner, value hypothesis and decision timetable.
Deployment, residency, egress, identity, retention, model and air-gap constraints.
Requirement ID, priority, vendor response, exception, evidence reference and owner.
Architecture, logs, policies, certifications, test outputs, references and dates.
Three volume scenarios, unit definitions, included capacity, indexation and implementation assumptions.
Environment, test set, attack cases, acceptance thresholds, data handling and deliverables.
Security, data, change, service, regulatory, pricing, IP, audit, exit and transition positions.
| Gate | Minimum pass condition |
|---|---|
| Deployment | The complete production architecture can run inside the required boundary. |
| Data use | Customer data is never used to train shared models or services without explicit, revocable instruction. |
| Identity | Every agent and privileged component has attributable, least-privilege identity and revocable secrets. |
| Action control | High-impact tools can be denied or approval-gated at runtime. |
| Audit | Every tested run can be reconstructed with actor, data, model, retrieval, tool, policy and approval events. |
| Portability | Agent definitions, prompts, evaluation assets, logs and source knowledge export in usable form. |
| Operations | The vendor supports the target network model, patch path, incident process and recovery objective. |
A failed gate is not repaired by a high total score. Record the exception owner, the compensating control and explicit risk acceptance — or eliminate the option.
| State | Meaning | Required follow-up |
|---|---|---|
| Comply | Available in the proposed configuration and contractually committed for go-live. | Evidence reference required |
| Comply with configuration | Requires documented configuration under customer control. | Configuration and test evidence |
| Partial | Meets part of the requirement; the limitation and workaround are explicit. | Gap, control owner and cost |
| Roadmap | Not available at the decision date. | Date, dependency and contractual remedy — score as absent |
| Exception | Vendor proposes a different control objective or boundary. | Residual risk and named buyer approval |
| No | Cannot meet the requirement. | Gate impact recorded |
Cap the rating a requirement can earn by the strongest evidence supplied for it. This single rule removes most presentation bias from an AI platform evaluation.
Written narrative or a demo statement
Current architecture, configuration, policy or sample evidence
Buyer observes the control in the proposed configuration
Comparable reference, independent assurance or repeatable production evidence
Traditional cloud questionnaires remain necessary but do not cover autonomous planning, indirect prompt injection, tool misuse, memory poisoning or cascading action. Convert each relevant agentic risk into an architecture requirement and a POC attack case.
| Attack surface | The question the vendor must answer with evidence |
|---|---|
| Goal and instruction integrity | How are system instructions separated from untrusted content, retrieved documents and tool results? |
| Tool misuse | How are tool schemas, destinations, parameters, rate, spend and action classes constrained at runtime? |
| Identity and privilege | Can every run prove which user, agent, workload credential and approval authorised each action? |
| Memory and context | How are memory writes validated, scoped, expired, reviewed and purged? |
| Inter-agent trust | How are messages authenticated, authorised, schema-validated and bounded across agents? |
| Failure propagation | What stops recursive loops, cascading retries, model fallback surprises and excessive spend? |
| Supply chain | How are models, connectors, tools, packages, prompts and update artefacts inventoried and verified? |
| Containment | Can the buyer revoke a tool, credential, model route or entire agent immediately — and prove it stopped? |
Regulation attaches to roles, use cases and sectors — not to a platform label. Require the vendor to distinguish the proposed system, the underlying models, and each party’s provider, deployer, controller, processor and ICT-service responsibilities.
| Regime | Procurement implication as of 9 August 2026 |
|---|---|
| EU AI Act — Article 50 | Transparency obligations for certain interactive and generative AI systems apply from 2 August 2026. Require interaction notices, marking and detection support, and deployer labelling controls where in scope. |
| EU AI Act — high-risk | After the 2026 AI Omnibus, Annex III high-risk rules apply from 2 December 2027, and high-risk AI embedded in Annex I products from 2 August 2028. Procure future evidence readiness now. |
| EU AI Act — GPAI | General-purpose AI provider obligations applied from 2 August 2025, with Commission enforcement powers from 2 August 2026. Identify model providers and downstream documentation dependencies. |
| GDPR | Existing controller and processor duties continue to apply: purpose, minimisation, lawful basis, security, data-subject rights, retention, subprocessors and international transfer controls. |
| DORA | Applicable since 17 January 2025 for in-scope financial entities. Article 30 requires specific ICT contract terms, and critical or important functions need documented, tested exit strategies. |
Do not ask a platform vendor to “certify our compliance.” Require the capabilities, documentation, cooperation, notice and evidence that let your organisation discharge its own obligations.
Notice configuration, marking or detection capability, disclosure workflow and retained proof.
Use-case assessment, provider/deployer statement, model-provider chain and change triggers.
Approval policy, reviewer information, override and stop capability, training and oversight events.
Data map, controller/processor terms, minimisation, retention, security, rights and subprocessor evidence.
ICT service register data, incident cooperation, audit rights, recovery evidence and a tested exit plan.
This is procurement guidance, not legal advice. Validate obligations, role classification and sector-specific requirements with qualified counsel and your control owners.
A proof-of-concept should resolve the highest-cost uncertainties in the target environment and produce evidence that survives the demo room.
List what the POC must prove and the artefact each test will produce.
Test the proposed network, identity, key, model and data-flow pattern — not a vendor-hosted substitute.
Keep part of the evaluation set from the vendor, including realistic ambiguity and low-quality source data.
Test hostile documents, over-privileged tools, malformed inputs, unavailable models and policy conflicts.
Record task quality, permission fidelity, action safety, latency, reliability, cost and trace completeness.
Set pass, conditional-pass and fail rules — including hard constraints — before the first vendor run.
| Test | Method | Acceptance principle |
|---|---|---|
| Workflow quality | Representative end-to-end cases plus a held-out set | Task-specific target agreed before the test |
| Permission fidelity | Authorised and unauthorised retrieval and tool cases | 100% of explicit deny cases denied and logged |
| Prompt injection | Direct and indirect malicious instructions | No prohibited disclosure or action; alerts and trace available |
| Approval gate | High-impact action requiring approval | No action before valid approval; approver and scope recorded |
| Boundary / egress | Network observation during representative runs | Zero unapproved external destinations or data classes |
| Trace reconstruction | An independent analyst reconstructs selected runs | All critical events and versions attributable |
| Failure and fallback | Unavailable model, connector timeout, malformed tool output | Bounded retry, safe state and an explicit degraded-mode signal |
| Kill switch | Revoke an agent, tool and credential mid-execution | New actions stop within the buyer-defined recovery target |
| Cost | Measured runs under realistic context and retry patterns | Observed units reconcile to the vendor estimate and invoice model |
| Portability | Export the POC workflow and evidence assets | Files complete, documented and usable without vendor access |
A hard-boundary failure — unauthorised data egress, an access-control breach, an untraceable action or a missing export — cannot be averaged away by quality or usability scores.
Use the same model for every vendor, publish the weights before final responses, and keep the solution rating separate from confidence in its evidence.
Deployment, residency, egress, key and operator control
Agent identity, least privilege, isolation, secrets and threat controls
Inventory, approval, autonomy, versioning, kill switch and exceptions
Complete traces, exports, reconstruction, regression and drift
Permission-aware retrieval, provenance, deletion and index control
Approved models, version pinning, route policy, fallback and BYOM
In-network connectors, schema controls, permissions and ownership
SLA, scaling, patch, recovery, incident and offline operation
Complete units, capacity transparency, indexation and stress-case cost
Usable export, dual run, deletion, assistance and tested continuity
Weighted score = Σ(weight × rating ÷ 5). A practical default is at least 70/100, every gate passed and no critical dimension below 3. Tune the thresholds to your risk appetite before issuing the RFP.
| Rating | Label | Anchor |
|---|---|---|
| 0 | Absent | Requirement is not met or not addressed. |
| 1 | Material gap | Conceptual capability, unsupported workaround or uncommitted roadmap. |
| 2 | Partial | Some coverage; significant manual process, limitation or compensating control. |
| 3 | Meets | Requirement is met in the proposed configuration with adequate evidence. |
| 4 | Strong | Exceeds the requirement; tested, operationally credible and easy to govern. |
| 5 | Differentiated | Material advantage with production-grade proof and a referenceable outcome. |
Move 5 points between control and feature dimensions — does the ranking change?
Reduce every untested critical claim to the document cap — does a vendor fall below threshold?
Use the stress workload and full support assumptions — is the recommendation still affordable?
Assume one compensating control fails or costs more — is residual risk still acceptable?
Loops, retries, long contexts, evaluation runs, re-indexing and human review can dominate the licence line. Compare the whole controlled service across three scenarios.
Approved adoption plan, normal context sizes, realistic retries and agreed service levels.
Peak concurrency, longer context, more tool calls, degraded dependencies and heavy evaluation or re-indexing.
Export, dual run, re-indexing, migration engineering, vendor assistance and retained access.
Licence + model inference + infrastructure + storage and indexing + integration + assurance + operations + change + support + exit rehearsal + contingency − verified process savings.
A low pilot estimate is not a production cost model. Reject pricing that cannot be reconciled to measured POC units and scaled across expected and stress workloads.
Attach the proposed configuration, the commitments and the evidence references to the contract, so the production service cannot quietly become a different risk.
Traditional SaaS release language is usually too broad for model-driven behaviour. Define which changes require notice, regression evidence, buyer approval or a right to remain on the prior version.
| Change class | Example | Contract treatment |
|---|---|---|
| Routine | Security patch with no material behaviour or data-flow change | Normal notice and release notes |
| Material | Model version, router, safety policy, retrieval, tool or logging change | Advance notice, impact statement and regression evidence |
| Boundary | New external service, location, subprocessor, telemetry or support path | Prior approval or a termination right |
| Emergency | Urgent security or safety response | Immediate containment, prompt notice, retrospective evidence |
| Customer-controlled | Agent, prompt, model or policy change initiated by the customer | Version, approval, test and rollback record |
| Exit stage | Acceptance evidence |
|---|---|
| Inventory | Agent definitions, prompts, policies, tools, models, RAG sources, indexes, evaluations, logs and dependencies. |
| Export | Open, documented formats with schema, version and completeness checks. |
| Reconstitute | Buyer or replacement supplier recreates a representative workflow without vendor control-plane access. |
| Dual run | Defined transition period, capacity, support and pre-agreed rates. |
| Delete | All vendor and subprocessor copies removed under a stated timeline; certificate and exceptions supplied. |
| Close | Revoke access, keys and identities; retain agreed evidence; resolve residual incidents and invoices. |
Test export and basic reconstitution during the POC, then attach the resulting package description to the exit schedule. Otherwise “exportable” remains undefined.
Transfer the decision record into an inventory, a control baseline, a monitoring plan and a reapproval calendar.
Business, technical, data, security, privacy, model and incident owners are named and trained.
Agent, purpose, users, risk tier, data, tools, models, version, owner and status are recorded.
Identity, policy, approval, logging, limits, kill switch, backup and recovery are tested.
Golden set, safety set, thresholds, baseline, drift signals and escalation are approved.
Monitoring, on-call, support, incident, change, access review and capacity processes are active.
Required user notices, disclosures, content marking or labelling and human escalation are implemented.
Budgets, cost allocation, alerts, caps, invoice reconciliation and forecast ownership are active.
Current export package, reconstitution test date, transition owner and deletion path are recorded.
| Cadence | Review |
|---|---|
| Continuous | Availability, errors, retries, policy denials, unusual tool use, spend, drift and safety signals. |
| Weekly | Workflow quality, escalations, failed cases, incidents, data freshness and cost variance. |
| Monthly | Access, agent inventory, model routes, exceptions, vendor issues and benefits trend. |
| Quarterly | Risk reclassification, control testing, red-team themes, capacity, TCO and exit-package currency. |
| On material change | Impact assessment, regression, security and privacy review, approval, communication and rollback plan. |
| Annually / renewal | Market test, references, service performance, contract protections and tested portability. |
Would we approve this system today, with its current models, data flows, actions, cost and evidence? If not, renewal should trigger remediation, re-scoping or exit — not automatic extension.
The 35-page PDF edition adds four appendices you can paste straight into a requirement matrix, a scoring sheet, a POC protocol and a pre-signature approval pack.
A minimum requirement library with IDs and a named minimum evidence artefact for each clause — governance, identity and action control, data and sovereignty, agentic security and supply chain, private RAG, model routing, audit and operations, commercial and exit.
Eligibility gate sheet, the ten weighted dimensions with rating, evidence cap and weighted-score columns, and a decision summary that records conditions precedent, residual risks, POC defects and the approval owner.
A worksheet you freeze before testing: workflow, target deployment, data classification, models and sources, hard-fail conditions, an eleven-test record and a POC decision with an issue and condition log.
The pre-signature gate across mandate, risk, boundary, RFP, scoring, POC, due diligence, TCO, contract, exit, go-live and approval — plus an approval record and open-conditions register.
Four layers: eligibility gates that eliminate non-viable vendors before scoring, a scored requirement matrix, an evidence schedule that names the artefact proving each critical requirement, and a POC protocol with acceptance thresholds agreed before testing. Requirements issued without a proof standard invite incomparable marketing answers.
Score written responses before any tailored demonstration, cap the rating a claim can earn by the strength of its evidence, and run the proof-of-concept in the target boundary against hostile, failure and recovery conditions. A polished narrative should never score above 2 out of 5.
An eligibility gate is a hard pass/fail condition — deployment boundary, data use, identity, action control, audit reconstruction, portability and operating model — that determines whether a vendor may be scored at all. A failed gate is never repaired by a high total score; it requires a named exception owner, a compensating control and explicit risk acceptance, or elimination.
Rate the action, not the interface. Autonomy is earned per action against impact, reversibility, data sensitivity, scale and observability — so one agent may recommend in one step, require approval in another and be prohibited from a third. If an agent can change a system of record, send externally, move money, alter access or affect rights, procure it as an action system rather than a chat tool.
Licence, model inference, infrastructure, storage and indexing, integration, assurance, operations, change, support, exit rehearsal and contingency, less verified process savings — modelled across expected, stress and exit scenarios. Loops, retries, long contexts, evaluation runs, re-indexing and human review routinely dominate the visible licence line.
As of 9 August 2026: EU AI Act Article 50 transparency duties apply from 2 August 2026, GPAI provider obligations have applied since 2 August 2025 with Commission enforcement powers from 2 August 2026, Annex III high-risk rules apply from 2 December 2027 after the AI Omnibus, GDPR controller and processor duties continue unchanged, and DORA has applied to in-scope financial entities since 17 January 2025 with prescriptive ICT contract terms and tested exit strategies.
Exportability is proven, not asserted: agent and workflow definitions with policies and version history, tool contracts and connector configuration, source inventory and ingestion configuration, golden sets and evaluation results, execution logs in a documented machine-readable format, model inventory and routing rules, and deployment runbooks. Test export and basic reconstitution during the POC, then attach the resulting package description to the exit schedule.
The guide is vendor-authored by VDF AI but deliberately structured as a buyer-controlled evaluation method: gates, evidence caps, weights and acceptance tests are ours to publish and yours to tune. Every requirement is written so any platform — including ours — can be failed by it.
The guide uses these primary and standards-body references as its regulatory and control baseline. Always check the latest official version before issuing an RFP or signing a contract.
Bring your gates, weights and evidence caps. We will walk the boundary declaration, the audit trace, the export package and the three-year cost model for the deployment you actually intend to run — and you can score the session.