All briefs
Brief 01/05 Financial Services Updated July 2026 8 min read

EXECUTIVE BRIEF · FINANCIAL SERVICES

Private AI for banking, without handing your data to a third party

Financial institutions are under pressure to deploy AI while satisfying DORA, GDPR, and internal risk controls. On-premises AI agents let you move fast on high-value workflows while keeping customer data, models, and audit trails inside your perimeter.

For CIOs, CISOs, Heads of Data, and Risk & Compliance leaders in banking and financial services.

The banking edition as a print-ready PDF — the compliance mapping, cost model, and 90-day plan in a format you can forward to risk and audit.

Mapped to DORAEU AI ActGDPR On-prem & sovereign cloud

The pressure

What is forcing the decision

Regulatory scrutiny is rising

DORA, the EU AI Act, and national supervisors expect demonstrable control over AI systems, third-party concentration risk, and operational resilience.

Data cannot leave the bank

Customer PII, transaction data, and market-sensitive information are subject to residency, secrecy, and contractual constraints that hosted AI cannot satisfy.

Cost and lock-in concerns

Per-token AI pricing is hard to forecast for high-volume workflows, and single-vendor model lock-in is itself a concentration risk.

ICT concentration risk under DORA

External inference providers become ICT third-party dependencies subject to concentration-risk assessment, contractual resilience clauses, and exit planning. On-prem deployment removes that dependency class entirely.

Why on-premises

The case for private deployment

On-Prem Private AI for Banking & Financial Services

On-premises deployment removes external inference as an ICT third-party dependency, keeps regulated data inside your control, and gives risk and audit teams a complete, inspectable trail for every AI action. With flat platform pricing — not per-token metering — high-volume banking workflows stay predictable. It is the most direct path to deploying AI in a way DORA and your supervisors will accept.

Compliance mapping

Mapped to your obligations

DORA

Removes third-party inference as an ICT concentration risk; full audit trail supports resilience testing.

EU AI Act

AI inventory, risk classification, and human-oversight controls for high-risk use cases.

GDPR

Data stays in-region and in-perimeter; supports minimization and erasure workflows.

Systems & data

Where the agents actually work

Core banking & payments
Read-only retrieval over account, transaction, and payment records so case context assembles itself instead of being copied between screens.
AML / case management
Alert queues, prior dispositions, and SAR narratives become searchable evidence rather than institutional memory held by senior analysts.
Policy, regulation & procedure libraries
Internal manuals, supervisory correspondence, and product terms indexed for cited answers — the single highest-volume query surface in most banks.
Credit files & counterparty data
Financials, covenants, and prior memos assembled into structured first drafts that analysts edit rather than compile.

First workflows

Where to start for fast payback

First workflows for deploying DORA-aligned private AI in banking and financial services.

  1. KYC / AML investigation support

    Agents assemble case context from internal systems, summarize alerts, and draft investigation notes — with every source and step logged.

  2. Regulatory and policy Q&A

    Private retrieval over internal policy, regulation, and procedure documents so staff get grounded answers with citations, never invented ones.

  3. Credit and risk memo drafting

    Agents compile structured data and documents into first-draft memos that analysts review, cutting cycle time without ceding judgment.

  4. Customer operations assist

    Frontline and back-office staff get AI assistance grounded in your own knowledge base, under RBAC, with no customer data leaving the bank.

The first 90 days

A phased path to production

  1. Days 0–30

    Contain the scope, prove the evidence trail

    Stand the platform up in a non-production zone against one document corpus — typically the internal policy and regulation library. The objective is not model output; it is showing risk and audit a complete, reproducible record of retrieval and generation for every answer.

    Exit criteriaModel-risk and second-line reviewers sign off on the logging and human-review design.

  2. Days 31–60

    Put it in front of first-line staff

    Roll policy and regulatory Q&A out to a defined population — compliance advisory, product, or branch support — under existing RBAC. Measure answer acceptance and the questions the corpus cannot answer, which is usually where the documentation debt sits.

    Exit criteriaSustained weekly usage from real staff, with a triaged list of retrieval gaps.

  3. Days 61–90

    Move to a regulated workflow

    Extend into KYC/AML case assembly or credit memo drafting, where the agent gathers evidence and the analyst retains the decision. Wire the audit pack into the format your supervisor already asks for during operational-resilience testing.

    Exit criteriaA production workflow with documented human-in-the-loop gates and an exportable audit pack.

The cost model

Why the economics work differently

Banking AI economics break in a specific place: the workflows worth automating are the high-frequency ones. A KYC refresh cycle or a policy Q&A rollout across thousands of staff generates query volumes where per-token pricing turns a capability decision into an unbounded OPEX exposure — one that finance functions increasingly treat as a material ICT cost risk in its own right.

Volume, not sophistication, drives the bill

Retrieval-heavy banking prompts carry large context windows. Cost scales with documents retrieved per question, which is precisely what improves answer quality.

GPU capacity is a fixed, depreciable asset

On-prem inference converts a variable cloud line item into infrastructure the bank already knows how to budget, finance, and amortise.

Exit costs are near zero

Models, embeddings, and indexes stay on your hardware, so switching model families is an engineering task rather than a renegotiation with a provider that holds your data.

Compare the two models in detail: committed flat pricing vs. pay-as-you-go.

Proof points

What similar organizations achieve

40–60% lower AI cost vs. metered cloud
−65% compliance reporting prep time
10× faster document processing

Proven in a Tier-1 retail bank

A European retail bank deployed on-prem KYC/AML investigation agents across 200 analysts. Case assembly time dropped while every retrieval source and model step remained in the audit pack supervisors requested under DORA operational-resilience testing.

Proven in an asset manager

An asset manager replaced hosted policy Q&A with private retrieval over internal compliance manuals. Analysts received cited answers only — no customer or position data left the perimeter — and model-risk teams retained full lineage for SR 11-7 documentation.

Objections

What buying committees push back on

"Our cloud provider already contractually guarantees data isolation."

Contractual isolation and architectural isolation fail differently. DORA reviews increasingly ask what happens when the provider is unavailable or the arrangement ends — an exit-planning question that on-prem inference answers structurally rather than in a clause.

"We do not have the GPU estate or the people to run this."

Most first deployments run on a single node sized for concurrent users rather than training. The operational surface is a containerised platform on hardware your infrastructure team already manages; model routing keeps smaller models on the high-volume paths so the estate stays modest.

"Open-weight models are behind the frontier."

For retrieval-grounded banking work — summarising an alert, citing a policy clause, drafting a memo section from supplied documents — the differentiator is retrieval quality and grounding discipline, not raw model reasoning. The architecture also lets you swap or add models as they improve, without moving your data.

Evaluation checklist

Questions to put to any vendor

  • Can the entire inference path — prompt, retrieval, embeddings, and logs — run with no outbound network connectivity?
  • Does the audit record reconstruct a specific historical answer, including the exact document versions retrieved?
  • Is pricing tied to platform capacity, or does it scale with tokens, runs, or seats as adoption grows?
  • How are model updates delivered and validated without re-opening a third-party risk assessment?
  • What happens to indexes, fine-tunes, and embeddings if the contract ends — who holds them, and in what format?

The full procurement version: Enterprise AI Agent RFP Checklist · On-Prem AI Reference Architecture

Questions

What leaders ask first

Does this satisfy DORA third-party risk requirements?

Running inference on-premises removes the external model provider as an ICT third-party dependency, which directly addresses concentration-risk concerns under DORA. Combined with full audit logging, it supports operational-resilience testing and reporting.

Can we keep customer data in-region?

Yes. All processing, retrieval, and model inference run inside your perimeter and can be pinned to a specific region or country, satisfying residency and banking-secrecy constraints.

How does on-prem AI compare to Azure OpenAI private endpoints for banking?

Private endpoints reduce public-internet exposure but the model service remains a cloud-managed third party. On-prem inference keeps the full data plane — prompts, embeddings, logs, and models — inside your ICT boundary, which is what DORA concentration-risk reviews increasingly expect for high-volume workflows.

Why is flat pricing important for financial services AI?

Banking workflows — KYC reviews, regulatory Q&A, customer operations — generate high query volumes. Per-token cloud pricing makes forecasting impossible and can itself become a material ICT cost risk. Flat on-prem platform pricing keeps spend predictable as adoption scales.

Which workflow should we start with?

Most banks begin with regulatory and policy Q&A or KYC/AML investigation support: high value, contained data scope, and an audit trail that satisfies risk teams before broader rollout.

What hardware does a bank need to run private AI?

A first production workflow typically runs on a single GPU server sized for concurrent users rather than for training, deployed into an existing data-centre zone. Model routing keeps high-volume, low-complexity requests on smaller models, so capacity grows with adoption instead of being provisioned for a worst case on day one.

The briefing

Thirty minutes, built around your constraints

For Financial Services, we walk your security, risk, and platform leads through the deployment model, the compliance position, and the first workflow worth funding. Three things we cover:

  1. Your DORA and model-risk position

    We map how on-prem inference changes your ICT third-party register and what your second line will still need to document.

  2. Reference architecture for your estate

    How the platform sits alongside core banking, your case-management stack, and existing RBAC — including the air-gapped variant where required.

  3. A costed first workflow

    One workflow chosen with you, with the capacity, integration effort, and review gates it actually needs.

Briefs for other regulated industries