Enterprise AI Economics

Buy, Lease, or Colocate: GPU Capacity for On-Prem AI

Deciding how much GPU capacity you need is the easy half. The harder question is where the hardware physically sits and how it is paid for — owned racks, colocation, leased hardware, or dedicated single-tenant hosting — and which of those still satisfies a sovereignty requirement.

The GPU sizing question is now well understood. Estimate the working set, add the KV cache, account for concurrency, and you get a defensible range — the arithmetic is laid out in our guide to estimating GPU requirements for local LLM workloads.

The question that follows it is answered far less carefully. Once you know roughly how many accelerators the platform needs, someone has to decide where they physically live, who owns them, and which budget line absorbs them. That decision has the longer shadow, because it is the one locked in by a contract, a facilities upgrade or a capital approval — and it determines whether the sovereignty claim in your architecture document survives contact with an auditor.

Four sourcing models, and what actually separates them

Own the hardware, run it in your own facility. The default reading of “on-premises”. Maximum control, and the only option where no third party has physical access to the machines. It assumes the facility can deliver the power and cooling envelope that modern accelerators demand — an assumption worth testing before it appears in a plan.

Own the hardware, run it in colocation. The enterprise buys and owns the servers; a colocation provider supplies the building, power, cooling and connectivity. Custody of data and administrative access remain with the enterprise. This is the pragmatic middle path for organisations whose own data centres were designed for 5 kW racks and cannot economically be retrofitted for AI density.

Lease or finance the hardware. Identical technically to owning, different on the balance sheet. Converts a large capital request into a predictable operating charge and shifts some refresh risk to the lessor. Useful when the capital approval process, rather than the money, is the binding constraint.

Dedicated single-tenant hosted capacity. Bare-metal accelerators reserved exclusively for one customer, operated by a provider. Faster to stand up than anything involving procurement of physical machines, but the provider holds administrative access to the platform layer unless the contract and the technical design say otherwise. This is where “private” and “dedicated” get used loosely, and where the diligence has to be sharpest.

A fifth route exists for burst training rather than production inference: shared national and European compute. The EuroHPC Joint Undertaking has been expanding access channels for industry alongside its AI Factories programme, and in July 2026 launched a call for up to seven AI Gigafactories across EU Member States. That capacity is well suited to occasional heavy jobs and poorly suited to a production workflow that must run at 09:00 every weekday.

The constraint that decides it is power, not silicon

Most enterprises approach this as a procurement question about accelerators. It is usually a facilities question about kilowatts.

Industry reporting through 2026 describes average rack densities rising sharply, with AI deployments commonly landing in the 50–70 kW band and the densest current systems far above that — figures that sit an order of magnitude beyond the general-purpose enterprise racks most corporate data halls were built around. Practitioner guidance now treats direct-to-chip liquid cooling as necessary above roughly 20 kW per rack, and purpose-built liquid infrastructure above roughly 50 kW.

Two consequences follow. First, the retrofit cost of an existing facility belongs in the sourcing comparison, not in a separate facilities budget where it stays invisible to the AI business case — a dependency that should appear in the total cost of ownership model from its first version. Second, availability is the real lead-time driver: reporting through 2026 has described colocation vacancy at record lows with much capacity under construction already precommitted, and enterprises reserving space 18 to 24 months ahead. A plan that assumes hardware delivery is the long pole will be wrong.

The finance question: useful life, and what happens after it

The depreciation schedule quietly determines whether owning looks cheaper than renting, and it is genuinely contested. Large cloud operators have converged on five- to six-year useful lives for AI servers; some specialist providers use four or five; well-publicised sceptics argue the economically useful life is closer to two or three given the pace of accelerator releases. The same hardware can look like a bargain or a write-off depending on an accounting assumption.

For an enterprise running production inference — rather than a provider competing to train frontier models — the sceptical case is weaker than it looks. Hardware that is no longer competitive for the largest workloads still serves smaller models, embedding and reranking work, and overnight batch jobs perfectly well. Capacity cascades down the stack rather than falling off a cliff, which is an argument for a platform that can route work across models of different sizes.

Two rules follow. Write the assumed useful life and second-phase use into the business case, so the number is inspectable rather than embedded. And compare owned against rented over the same horizon — a five-year ownership case set against a twelve-month rental quote is not a comparison.

The sovereignty test each option has to pass

Sourcing models are argued in terms of cost and lead time, then quietly assumed to be equivalent on control. They are not. Four questions separate them, and all four should be answered in writing before the commercial discussion:

  1. Who holds physical access? Anyone who can touch the machine is in scope for the security model, including facility staff under an escort regime.
  2. Who holds administrative credentials to the platform and the hypervisor? This most often distinguishes genuinely dedicated capacity from a well-marketed shared service.
  3. Where do the derived artefacts live? Prompts, embeddings, vector indexes and audit trails carry as much sensitive content as the source documents — the point developed in our analysis of data sovereignty risks for regulated industries.
  4. What is the exit? Does moving the workload require re-engineering the platform, or only re-hosting it?

A clause promising that data will not leave a region is a commitment; holding the only set of administrative keys is a control. Regulated buyers increasingly need the second and get asked to demonstrate it — the same distinction that separates data sovereignty from data residency.

A decision sequence that avoids the common traps

Work through these in order. Reversing them is how organisations end up with accelerators they cannot power.

  1. Classify the workload shape — steady production inference, bursty experimentation, or periodic training. Each has a different economic answer, and most enterprises have at least two.
  2. Test the facility envelope honestly: power per rack, cooling, floor loading, and the cost and timeline of closing any gap. Where retrofit is expensive, colocation with owned hardware is usually the shorter path.
  3. Fix the control requirements by answering the four questions above, and treat the answers as requirements rather than criteria to be traded away later.
  4. Model capital and operating cases over one horizon, including facility cost, network, platform staff, support and the refresh assumption — not the accelerator price list.
  5. Design for portability regardless of the answer.

What must stay constant across every option

The most expensive mistake here is letting the sourcing decision dictate the platform architecture. Infrastructure sourcing changes on a three-to-five-year rhythm as facilities, prices and hardware generations shift; the AI platform — model registry, retrieval pipeline, agent definitions, governance and audit — should outlive several of those cycles.

So the platform layer needs to behave identically on owned racks in a corporate hall, on owned racks in a colocation cage, and on dedicated hosted hardware. If moving the workload between them requires rebuilding workflows, the sourcing choice has silently become a lock-in decision — the outcome the vendor lock-in evaluation is meant to prevent.

How VDF AI fits the sourcing decision

VDF AI deploys into infrastructure the customer controls, and does not care which of the four sourcing models it came from: the same platform runs on owned hardware in a corporate data centre, on customer-owned hardware in colocation, or on dedicated single-tenant capacity, with models, documents, embeddings, logs and audit records staying inside that boundary in every case.

Because licensing is capacity-based rather than metered per token, the infrastructure and software decisions stay independent — changing where the hardware sits does not reprice the platform. Multiple local models of different sizes can be registered and routed against, so hardware past its first-choice role keeps carrying useful work. And because deployment, governance and audit behave identically across the options, sourcing stays what it should be: a facilities and finance question, reversible on its own timetable.

Further reading

Sources


Working out where your AI hardware should live? See how VDF AI deploys into infrastructure you control, or book a demo.

Frequently asked questions

Does colocation still count as on-premises AI?

For most regulatory and procurement purposes, what matters is not the postcode of the building but who controls the hardware, the data on it, and administrative access to both. Colocated hardware that the enterprise owns, patches and holds the only credentials for is a substantially different arrangement from renting time on shared infrastructure, even though neither sits in a company-owned building. The distinction worth documenting is single-tenancy and control, not the word on-premises — because that is the distinction an examiner or an auditor will actually probe.

Should an enterprise buy GPUs or rent them for AI inference?

It depends on how steady the workload is. Inference for production business workflows tends to be predictable and continuous, which favours owned or leased capacity where the cost per unit of throughput falls as utilisation rises. Bursty, exploratory or one-off training work has the opposite shape and is usually cheaper to rent. Many enterprises end up with both: owned capacity sized to the steady production baseline, plus a rented or shared-facility route for occasional heavy jobs.

How long should GPU hardware be depreciated over?

There is no single answer, and the range in use is wide. Large cloud operators have generally settled on five- to six-year useful lives for AI servers, while some smaller specialist providers use four or five years, and prominent critics argue the economically useful life is shorter. For an enterprise inference platform the practical point is that hardware does not become worthless when it stops being state of the art — it moves down the stack to smaller models, batch work and non-latency-sensitive tasks. Pick a schedule your finance function can defend and state the assumption explicitly in the business case.

What is the biggest planning mistake in sourcing AI infrastructure?

Underestimating lead time on power and cooling rather than on the accelerators themselves. High-density AI racks need far more power per rack and, above roughly 20 kW, direct-to-chip liquid cooling — which many existing enterprise facilities were never built to provide. Industry reporting through 2026 describes very tight colocation availability and enterprises securing space 18 to 24 months ahead of deployment. The GPUs are frequently the fastest part of the order to arrive.

Filed under
on-premises AIAI platform TCOAI procuremententerprise AI investmentAI infrastructure
AI Cost & Energy

Calculate your AI infrastructure savings

Model the cost and energy impact of running AI on-prem versus cloud-only — then see the benchmark data behind the numbers.

Keep reading