The GPU sizing question is now well understood. Estimate the working set, add the KV cache, account for concurrency, and you get a defensible range — the arithmetic is laid out in our guide to estimating GPU requirements for local LLM workloads.
The question that follows it is answered far less carefully. Once you know roughly how many accelerators the platform needs, someone has to decide where they physically live, who owns them, and which budget line absorbs them. That decision has the longer shadow, because it is the one locked in by a contract, a facilities upgrade or a capital approval — and it determines whether the sovereignty claim in your architecture document survives contact with an auditor.
Four sourcing models, and what actually separates them
Own the hardware, run it in your own facility. The default reading of “on-premises”. Maximum control, and the only option where no third party has physical access to the machines. It assumes the facility can deliver the power and cooling envelope that modern accelerators demand — an assumption worth testing before it appears in a plan.
Own the hardware, run it in colocation. The enterprise buys and owns the servers; a colocation provider supplies the building, power, cooling and connectivity. Custody of data and administrative access remain with the enterprise. This is the pragmatic middle path for organisations whose own data centres were designed for 5 kW racks and cannot economically be retrofitted for AI density.
Lease or finance the hardware. Identical technically to owning, different on the balance sheet. Converts a large capital request into a predictable operating charge and shifts some refresh risk to the lessor. Useful when the capital approval process, rather than the money, is the binding constraint.
Dedicated single-tenant hosted capacity. Bare-metal accelerators reserved exclusively for one customer, operated by a provider. Faster to stand up than anything involving procurement of physical machines, but the provider holds administrative access to the platform layer unless the contract and the technical design say otherwise. This is where “private” and “dedicated” get used loosely, and where the diligence has to be sharpest.
A fifth route exists for burst training rather than production inference: shared national and European compute. The EuroHPC Joint Undertaking has been expanding access channels for industry alongside its AI Factories programme, and in July 2026 launched a call for up to seven AI Gigafactories across EU Member States. That capacity is well suited to occasional heavy jobs and poorly suited to a production workflow that must run at 09:00 every weekday.
The constraint that decides it is power, not silicon
Most enterprises approach this as a procurement question about accelerators. It is usually a facilities question about kilowatts.
Industry reporting through 2026 describes average rack densities rising sharply, with AI deployments commonly landing in the 50–70 kW band and the densest current systems far above that — figures that sit an order of magnitude beyond the general-purpose enterprise racks most corporate data halls were built around. Practitioner guidance now treats direct-to-chip liquid cooling as necessary above roughly 20 kW per rack, and purpose-built liquid infrastructure above roughly 50 kW.
Two consequences follow. First, the retrofit cost of an existing facility belongs in the sourcing comparison, not in a separate facilities budget where it stays invisible to the AI business case — a dependency that should appear in the total cost of ownership model from its first version. Second, availability is the real lead-time driver: reporting through 2026 has described colocation vacancy at record lows with much capacity under construction already precommitted, and enterprises reserving space 18 to 24 months ahead. A plan that assumes hardware delivery is the long pole will be wrong.
The finance question: useful life, and what happens after it
The depreciation schedule quietly determines whether owning looks cheaper than renting, and it is genuinely contested. Large cloud operators have converged on five- to six-year useful lives for AI servers; some specialist providers use four or five; well-publicised sceptics argue the economically useful life is closer to two or three given the pace of accelerator releases. The same hardware can look like a bargain or a write-off depending on an accounting assumption.
For an enterprise running production inference — rather than a provider competing to train frontier models — the sceptical case is weaker than it looks. Hardware that is no longer competitive for the largest workloads still serves smaller models, embedding and reranking work, and overnight batch jobs perfectly well. Capacity cascades down the stack rather than falling off a cliff, which is an argument for a platform that can route work across models of different sizes.
Two rules follow. Write the assumed useful life and second-phase use into the business case, so the number is inspectable rather than embedded. And compare owned against rented over the same horizon — a five-year ownership case set against a twelve-month rental quote is not a comparison.
The sovereignty test each option has to pass
Sourcing models are argued in terms of cost and lead time, then quietly assumed to be equivalent on control. They are not. Four questions separate them, and all four should be answered in writing before the commercial discussion:
- Who holds physical access? Anyone who can touch the machine is in scope for the security model, including facility staff under an escort regime.
- Who holds administrative credentials to the platform and the hypervisor? This most often distinguishes genuinely dedicated capacity from a well-marketed shared service.
- Where do the derived artefacts live? Prompts, embeddings, vector indexes and audit trails carry as much sensitive content as the source documents — the point developed in our analysis of data sovereignty risks for regulated industries.
- What is the exit? Does moving the workload require re-engineering the platform, or only re-hosting it?
A clause promising that data will not leave a region is a commitment; holding the only set of administrative keys is a control. Regulated buyers increasingly need the second and get asked to demonstrate it — the same distinction that separates data sovereignty from data residency.
A decision sequence that avoids the common traps
Work through these in order. Reversing them is how organisations end up with accelerators they cannot power.
- Classify the workload shape — steady production inference, bursty experimentation, or periodic training. Each has a different economic answer, and most enterprises have at least two.
- Test the facility envelope honestly: power per rack, cooling, floor loading, and the cost and timeline of closing any gap. Where retrofit is expensive, colocation with owned hardware is usually the shorter path.
- Fix the control requirements by answering the four questions above, and treat the answers as requirements rather than criteria to be traded away later.
- Model capital and operating cases over one horizon, including facility cost, network, platform staff, support and the refresh assumption — not the accelerator price list.
- Design for portability regardless of the answer.
What must stay constant across every option
The most expensive mistake here is letting the sourcing decision dictate the platform architecture. Infrastructure sourcing changes on a three-to-five-year rhythm as facilities, prices and hardware generations shift; the AI platform — model registry, retrieval pipeline, agent definitions, governance and audit — should outlive several of those cycles.
So the platform layer needs to behave identically on owned racks in a corporate hall, on owned racks in a colocation cage, and on dedicated hosted hardware. If moving the workload between them requires rebuilding workflows, the sourcing choice has silently become a lock-in decision — the outcome the vendor lock-in evaluation is meant to prevent.
How VDF AI fits the sourcing decision
VDF AI deploys into infrastructure the customer controls, and does not care which of the four sourcing models it came from: the same platform runs on owned hardware in a corporate data centre, on customer-owned hardware in colocation, or on dedicated single-tenant capacity, with models, documents, embeddings, logs and audit records staying inside that boundary in every case.
Because licensing is capacity-based rather than metered per token, the infrastructure and software decisions stay independent — changing where the hardware sits does not reprice the platform. Multiple local models of different sizes can be registered and routed against, so hardware past its first-choice role keeps carrying useful work. And because deployment, governance and audit behave identically across the options, sourcing stays what it should be: a facilities and finance question, reversible on its own timetable.
Further reading
- How to Estimate GPU Requirements for Local LLM Workloads
- On-Premise AI Platform Cost and TCO Guide
- Chargeback and Showback for a Shared On-Prem AI Platform
- Data Sovereignty vs Data Residency in AI Procurement
- How to Evaluate Vendor Lock-In in Enterprise AI Platforms
Sources
- EuroHPC JU — The EuroHPC Joint Undertaking launches the AI Gigafactories Call
- Schneider Electric — Planning liquid-cooled AI data centers around grid and power constraints
- Inflect — Data Center Colocation Trends 2026: Power, AI Density, and the New Rules of Colocation
- CNBC — The question everyone in AI is asking: how long before a GPU depreciates?
Working out where your AI hardware should live? See how VDF AI deploys into infrastructure you control, or book a demo.