Most enterprise AI platform selections are decided by demo. A vendor arrives with a scenario they have run a hundred times, on data they curated, in an environment they control, and the room comes away impressed. Three vendors do the same thing, the scores are close, and the decision quietly gets made on rapport, brand or price.
That process reliably picks the best demo. It does not reliably pick the platform that will still be running the workload in three years, which is a different question, and one the demo was never designed to answer.
A bake-off — a competitive proof of value in which every vendor builds the same thing, on the same data, against the same acceptance criteria — answers it far better. It is more work to organise, and the work sits with the buyer rather than the vendor. That is precisely why it produces a decision that survives contact with procurement, risk and the board.
Pick a workload that is boring and real
The instinct is to choose the most exciting use case. Resist it. The workload that discriminates between platforms is one that is representative of your operating reality rather than your ambition.
Good candidates share four properties. They touch at least two internal systems, because integration is where platforms diverge and slideware converges. They involve documents or data with real access restrictions, so permission handling is exercised rather than described. They have an outcome someone can judge as right or wrong without a workshop. And they matter enough that the business will assign a subject-matter expert to the evaluation — an evaluation without a domain expert scoring the output is a technical exercise, not a procurement decision.
Deliberately avoid the workload where a single model’s raw capability dominates the result. If a task can be done by prompting any competent model, it tells you about the model, not the platform. What you are buying is the layer around the model: retrieval, tool access, orchestration, approvals, audit, deployment and operations. Choose a task that exercises that layer.
Fix the data set before the vendors arrive
The most common failure in a bake-off is that each vendor ends up with slightly different data — because access took too long to arrange, because one team was given a cleaner extract, because a late redaction pass removed the hard cases.
Assemble a frozen evaluation set in advance: a document corpus, a set of representative inputs and, critically, a held-back set the vendors never see before the final run. Include the awkward examples on purpose. Scanned documents, tables inside PDFs, near-duplicate policy versions, records with restricted access, non-English content if your organisation has it. The retrieval failure modes that matter in production are almost never visible on a clean corpus.
Write the expected answers before the evaluation, not after. Retrofitting acceptance criteria to what the systems produced is how a bake-off turns into a rationalisation.
Score four things, not forty
Long weighted scorecards feel rigorous and behave badly: with enough criteria, every vendor scores about the same, and the tie is broken by whoever filled in the questionnaire most enthusiastically. Four dimensions carry most of the decision.
Task outcome. Did the workload produce correct, complete, usable results on the held-back set, judged by the domain expert rather than the project team? Record the failures as well as the rate — a platform that fails visibly and traceably is safer than one that fails plausibly.
Evidence. For each result, can you reconstruct what happened: which model version, which retrieved sources, which tools were called with which arguments, who approved what and when? Ask for this as a produced artefact during the pilot, not a roadmap commitment. This is the same observability trail your auditors will ask for later, and it is far cheaper to discover its absence now.
Change cost. Midway through the evaluation, change the requirement. Add a validation rule, restrict a data source to one department, insert a human approval step before a system write. The time and skill each platform needs to absorb that change predicts your next three years better than any feature comparison, because the requirement will change.
Operational fit. Who runs it, on what hardware, with what upgrade path, and what happens when a model or a dependency needs replacing? Separating development, test and production environments is a hard requirement in most regulated estates and an afterthought in many platforms.
Make the contract questions part of the exercise
In regulated sectors, the evaluation and the contract are not separate workstreams. For financial entities in the EU, DORA sets out what an ICT services contract must contain — Article 30(2) covers matters such as service descriptions, the locations where services are provided and data is processed, and provisions on data availability, integrity and confidentiality, with Article 30(3) adding requirements for services supporting critical or important functions, including full service level descriptions, audit rights and exit strategies.
Those clauses are far easier to negotiate when the pilot has already produced the evidence behind them. Ask, during the bake-off, where processing and storage would occur in production, what the vendor’s subcontracting chain looks like, what SLAs are measurable rather than aspirational, and what an exit actually yields in files. Then test the exit: export everything the pilot produced and see what is portable. That is the difference between an assessed lock-in risk and a stated one.
The AI Act sequencing matters here too. The Digital Omnibus on AI, in force since 27 July 2026, deferred the high-risk obligations for Annex III systems to 2 December 2027 and for AI embedded in regulated products to 2 August 2028, while the Article 50 transparency obligations applied from 2 August 2026. The deferral is preparation time, not relief: documentation, logging and human oversight still have to be evidenced, and a platform that cannot demonstrate them in a controlled pilot is unlikely to produce them under supervisory pressure.
Run it inside your boundary
If the eventual deployment is on-premises or in a private environment, the bake-off should be too. A pilot run on a vendor’s hosted sandbox tests a product you are not buying, and it moves real data outside the boundary to do so — which often means the pilot corpus gets sanitised, which removes exactly the cases that discriminate.
Running each candidate inside your own environment costs more coordination and reveals more: how the platform behaves on your GPUs, against your identity provider, through your network controls, with your change process. Several evaluations end at this step, because a platform that is straightforward in a vendor tenancy turns out to need egress, managed services or elevated privileges that your architecture will not grant.
What a good outcome looks like
A finished bake-off should leave you with four things: a working implementation of a real workload, a scored comparison a domain expert stands behind, an evidence pack that risk and audit have already reviewed, and a short list of the things that went wrong — because those are what production will amplify.
If the exercise instead produced a slide saying all three vendors are broadly capable, the workload was too easy or the acceptance criteria were written too late. That result is worth knowing before the contract, not after.
How VDF AI approaches evaluations
VDF AI is designed to be evaluated the way it is deployed: inside the customer’s own data centre or private environment, on the customer’s hardware, against the customer’s data. Workflow definitions, retrieval configuration, approval steps and audit logs are artefacts you can inspect and export during the pilot, so the evidence pack the evaluation produces is the same evidence pack that supports the contract and the later governance review.
The change-cost test tends to be the informative one. Adding a validation rule, restricting a source to one department or inserting a human approval gate is configuration in the platform rather than a development cycle — which is the property that determines whether a pilot becomes a production system or a proof that nobody maintains.
Further reading
- Enterprise AI Procurement Checklist for Regulated Organizations
- How to Compare AI Agent Platforms Beyond Feature Lists
- How to Evaluate Vendor Lock-In in Enterprise AI Platforms
- AI Pilot Costs vs Production AI Platform Costs
- EU AI Act Enforcement Timeline and Deployer Readiness
Sources
- Regulation (EU) 2022/2554 (DORA) — Article 30, Key contractual provisions
- European Commission — Transparency obligations under Article 50 of the AI Act
- Gibson Dunn — EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes
- NIST — AI Risk Management Framework (AI RMF 1.0)
Planning a platform evaluation? See how VDF AI runs governed agent workloads inside your own environment, or book a demo.