AI Infrastructure

On-Premise Document AI: Azure, Google, SAP, Mistral and Open-Source Options Compared

Which document AI services run on premise in October 2026, what each one requires, and how open models such as PaddleOCR, Docling, olmOCR and DeepSeek-OCR compare on licence, layout and tables, handwriting, languages and hardware, with a selection checklist for intelligent document processing.

On-premise document AI is OCR, layout analysis and field extraction running on infrastructure you control. Of the cloud services compared here, Azure Document Intelligence is the one you can run in your own containers, though a fully disconnected setup needs Microsoft's approval and a commitment plan. Google and SAP run Document AI as cloud services, Mistral self-hosts OCR for selected enterprises, and open models cover the rest.

What on-premise document AI has to do

Intelligent document processing (IDP) is a pipeline. On premise, each stage needs a component you run yourself:

  1. Intake and classification. Decide what each file is (invoice, ID, contract, claim form) and whether it already carries a usable text layer.
  2. Text recognition. OCR for scans and photographs, direct text extraction for digitally produced PDFs.
  3. Layout and tables. Reading order, headings, table cells, checkboxes and key-value pairs.
  4. Field extraction. Named values for each document class, from templates, trained extractors or a language model.
  5. Validation and review. Business rules, confidence thresholds and a human queue for doubtful fields.
  6. Hand-off. Structured output into an ERP, a case system or a search index, linked back to the page image.

Cloud IDP services bundle most of these stages behind one API. Self-hosting means choosing a component per stage: more integration work, in exchange for control over every model, log and retention rule. The retrieval side of the same problem, including chunking, provenance and confidence routing for private RAG, is covered in our scanned-documents deep dive. If invoices are the first document class, our post on invoice exception agents picks up after extraction.

Cloud document AI services with an on-premise route

ServiceOn-premise option (verified October 2026)Conditions
Azure Document IntelligenceDocker containers, connected or disconnectedConnected mode meters usage through Azure; disconnected mode needs approval, a commitment plan and an expiring license file
Google Document AINone in the Document AI documentationProcessors run in Google Cloud locations; Google Distributed Cloud air-gapped offers a separate OCR API
SAP Document AINone foundManaged service on SAP BTP that reaches on-premise SAP systems through SAP Cloud Connector
Mistral OCR 4 and 4.1Self-hosted single container for enterprise customersProprietary model; API at $4 per 1,000 pages, $2 with the Batch API (list, verified October 2026)

Azure Document Intelligence containers

Microsoft publishes Azure Document Intelligence in Foundry Tools as Docker images in the Microsoft Container Registry. Which models you can run depends on the API version:

API versionModels available as containers
v4.0 (2024-11-30)Read, Layout
v3.1 (2023-07-31)Read, Layout, ID Document, Receipt, Invoice
v3.0 (2022-08-31)Read, Layout, General Document, Business Card, Custom (template)

The newest containers therefore cover OCR and layout only. Prebuilt invoice, receipt and ID extraction means running v3.1 images, custom extraction means v3.0 template models, and Microsoft’s documentation already marks v3.0 as retiring. Several containers also depend on another one: Invoice needs Layout, Receipt needs Read or Layout, and ID Document and Business Card need Read.

Microsoft’s sizing table for disconnected use asks for 8 cores per container, with minimum memory from 8 GB for ID Document to 16 GB for Layout, Invoice and Custom Template, and 24 GB recommended across the board. Connected containers report usage to Azure for billing, and Microsoft states that Azure containers do not send customer data such as the analysed image or text. Disconnected use starts with a request form and a commitment-tier resource. The container then downloads a license file while still connected, and that file carries an expiry date. On handwriting, the v4.0 Read and Layout models list 12 languages: English, Chinese Simplified, French, German, Italian, Thai, Japanese, Korean, Portuguese, Spanish, Russian and Arabic.

Google Document AI and Google Distributed Cloud

Google describes Document AI as a platform that turns unstructured document data into structured data. Processors are created in the us or eu multi-regions or in one of seven single regions with limited support, and the Document AI documentation lists no container or on-premises edition.

For disconnected sites, Google Distributed Cloud air-gapped brings Vertex AI into on-premises environments. Its three pre-trained APIs are OCR, Speech-to-Text and Translation. The OCR service detects printed and handwritten text in JPG, PNG, PDF and TIFF files, but Document AI’s processors are not on that list, so treat it as a text-recognition API inside Google’s air-gapped platform.

SAP Document AI

SAP’s reference architecture describes Document AI as a managed service on SAP Business Technology Platform that combines OCR, pretrained and customer-trained extraction models and LLM-based reasoning. SAP Cloud Connector gives it a path into on-premise systems such as SAP ECC and SAP S/4HANA, so extracted values can land in an on-premise ERP while the documents themselves are processed on SAP BTP.

Mistral OCR

Mistral announced OCR 4 on 23 June 2026, and OCR 4.1 became generally available on 16 July 2026. Both are proprietary models in Mistral’s Premier tier, priced as shown in the table above. Mistral says OCR 4 is compact enough to deploy in a single container on your own infrastructure and offers that self-managed deployment to enterprise customers. OCR 4 returns bounding boxes, block types such as tables and signatures, and per-page and per-word confidence scores, and Mistral claims support for 170 languages across 10 language groups.

Open OCR and document parsing models compared

These projects run wherever you have the hardware. Licences, features and requirements come from each project’s repository or model card (verified October 2026). “Not stated” means the project makes no claim either way, so test it yourself.

ProjectApproachLicenceLayout and tablesHandwritingLanguagesHardware
TesseractLSTM engine focused on line recognitionApache 2.0Text with positions; hOCR, ALTO, PAGE, TSV and searchable PDF outputFAQ: will not work very well; designed for printed text100+ out of the boxCPU
PaddleOCRPP-OCRv5 and PP-OCRv6 pipelines, PP-StructureV3 parsing, the 0.9B PaddleOCR-VL modelApache 2.0PP-StructureV3 returns table cell and text coordinates as Markdown or JSONPaddleOCR-VL described as suited to handwritten and historical documentsPP-OCRv5: 109; PP-OCRv6: 50 in one model; PaddleOCR-VL-1.5: 111CPU, GPU, XPU and NPU; ONNX Runtime, OpenVINO and TensorRT
docTRTwo-stage text detection and recognition on PyTorchApache 2.0Optional layout detection (titles, text, tables, page headers and footers) and a key information extraction predictorNot statedCharacter vocabularies defined for dozens of languagesPyTorch
DoclingConverter with layout and table-structure models, plus an optional VLM pipelineMIT; Granite-Docling-258M weights Apache 2.0Page layout, reading order, table structure, code and formulasNot statedGranite-Docling: English, with experimental Japanese, Arabic and ChineseDocumented for local and air-gapped execution
MarkerConverts PDF, images and Office files to Markdown, JSON, HTML or chunksCode Apache 2.0; weights under a modified OpenRAIL-M licenceTable converter; optional LLM mode merges tables across pagesNot statedClaims all languagesGPU, CPU or Apple MPS
olmOCRToolkit around a 7B VLM fine-tuned from Qwen2.5-VLApache 2.0Tables, equations and multi-column reading order; strips headers and footersListed as supportedNot statedNVIDIA GPU with at least 12 GB; 30 GB of free disk
DeepSeek-OCR and DeepSeek-OCR-2VLMs of about 3B parameters, steered by promptsMIT (OCR); Apache 2.0 (OCR-2)Prompt modes for document-to-Markdown with layout grounding and for parsing figuresNot statedNot statedNVIDIA GPU; supported in upstream vLLM

A few notes the table cannot carry:

  • Tesseract suits clean printed scans, and its maintainers say plainly that better results often depend on improving image quality first.
  • PaddleOCR moves quickly. Release 3.7.0 shipped in June 2026, and PP-OCRv6 recognises its 50 languages with a single unified model, so multilingual pages need no model switching.
  • Docling is the parser the scanned-documents post examines in depth; for this comparison, what matters is its MIT licence and its explicit support for air-gapped environments.
  • olmOCR release notes for version 0.3.0 mention fixing hallucinations on blank documents, a reminder that VLM-based OCR fails differently from classic engines.
  • DeepSeek-OCR was released on 20 October 2025 and added upstream vLLM support three days later; DeepSeek-OCR-2 followed in January 2026.

Licence and data-flow traps to check before a pilot

  • Marker’s model weights are not free for everyone. The modified OpenRAIL-M licence covers research, personal use and startups under $5M in funding or revenue; larger commercial users need a licence from Datalab.
  • Marker’s LLM mode calls Gemini by default. The optional —use_llm flag improves tables and forms, and out of the box it uses a Google Gemini model. On premise, point it at a local model served through Ollama or leave it off.
  • olmOCR can talk to hosted endpoints. Its README lists external inference providers it has tested; make sure the server flag points at your own vLLM instance.
  • Docling and Granite-Docling fetch models on first use. The Granite-Docling card notes that Docling downloads the model automatically, so mirror the weights before moving into an air-gapped network.
  • Azure’s disconnected licence expires. Plan for the commitment term and fresh license files when you pull new container images, because Microsoft warns an old file may fail to start a new image.
  • Mistral’s self-hosted OCR is a commercial agreement. Budget for negotiation time, since there are no published weights to evaluate first.

Hardware for document AI on premise

Document AI splits into three hardware tiers:

  • CPU tier. Tesseract runs on CPUs, and PaddleOCR supports Intel CPUs with OpenVINO acceleration, which suits high-volume printed scans with no GPU budget.
  • Container tier. Azure’s containers run on x64 CPU hosts: plan 8 cores and up to 24 GB of memory per container, plus any supporting container the model needs.
  • GPU tier. Vision-language models need accelerators. olmOCR asks for an NVIDIA GPU with at least 12 GB and was tested on RTX 4090, L40S, A100 and H100. DeepSeek-OCR’s PDF throughput figure comes from an A100 40 GB, and its roughly 3 billion parameters ship in BF16.

For GPU sizing, apply the weights and KV-cache arithmetic from our GPU estimation guide to the VLM you pick. Backfilling an archive is a long batch job that can starve interactive users, so give it its own queue with GPU admission control.

Accuracy risks with vision-language OCR

Classic OCR engines tend to fail visibly, with garbled characters on a poor scan. Vision-language models can fail fluently, producing clean text that is not on the page, as olmOCR’s fix for blank documents shows. Three habits contain both kinds of error:

  1. Ask for confidence. Mistral OCR 4 returns per-page and per-word confidence scores. Where an open model returns none, use proxies such as agreement between two engines, and route disagreements to review.
  2. Keep the page image beside every value. Reviewers should see the source region next to the extracted field, and the stored record should carry document, page and region identifiers.
  3. Test on your worst documents. Faded photocopies, rotated scans, stamps, handwriting and mixed languages expose differences that clean benchmark pages hide.

A selection checklist for on-premise IDP

  1. Inventory document classes and volumes. Pages per day, peak backfill size and which classes carry personal or regulated data.
  2. Split digital from scanned at intake. Digitally produced files rarely need OCR at all.
  3. Set a processing boundary per class. On premise only, disconnected, or an approved cloud region.
  4. Read the licences for code and weights. Look for revenue thresholds, non-commercial clauses and commitment terms.
  5. Match language and handwriting coverage. Compare each class against the vendor’s or project’s own published lists, not marketing summaries.
  6. Demand tables as structure. Cell coordinates and spans, not flattened text.
  7. Require confidence or build a proxy. Then set thresholds per field, not per document.
  8. Design human review into the flow. Side-by-side page image, field value and an audit trail of corrections.
  9. Keep provenance. Document ID, version, page and region for every extracted value.
  10. Size hardware per tier. Separate CPU, container and GPU workloads, and isolate batch backfills.
  11. Plan updates. Mirror models, pin versions, re-run your evaluation set, and track licence expiry for vendor containers.
  12. Log every run. Which document, which model version, which reviewer and what changed.

How VDF AI fits

VDF AI handles document work as tools that agents call inside your environment. The document parser tool turns files of most formats into clean text and structure, the table extraction tool recovers headers, rows and cells from documents, PDFs and images, and the OCR tool reads text from scans and photos. With the Tesseract engine selected, OCR runs entirely inside your perimeter with multi-language support and audit logging; a vision-model engine is available for difficult images.

For analysts, the report analysis product accepts PDF, Word, Excel, CSV and PowerPoint files as well as JPG, PNG and TIFF images read with OCR, and answers questions about them in a chat interface.

Sources

Frequently asked questions

Can Azure Document Intelligence run on premise?

Yes, as Docker containers on your own x64 hosts. In October 2026 the v4.0 containers cover the Read and Layout models, the v3.1 containers add ID document, receipt and invoice, and the v3.0 containers include general document, business card and custom template models. Connected containers report usage to Azure for billing, and Microsoft says they do not send the documents themselves. Running fully disconnected requires an approved application, a commitment-tier resource and a license file that expires, so plan renewals into operations.

Is Google Document AI available on premise?

Not as Document AI. Google runs Document AI processors in Google Cloud, in the us or eu multi-regions or a short list of single regions, and the Document AI documentation lists no on-premises edition. Google Distributed Cloud air-gapped, which brings Vertex AI into on-premises environments, includes an OCR API among its three pre-trained APIs. It detects printed and handwritten text in images, PDF and TIFF files, but it is an OCR service, not the Document AI processor catalogue.

Can Mistral OCR be self-hosted?

Only by arrangement with Mistral. OCR 4, announced in June 2026, is a proprietary model sold through Mistral's API at $4 per 1,000 pages, or $2 through the Batch API (list, verified October 2026). Mistral says the model is compact enough to run in a single container on your own infrastructure and offers that deployment to enterprise customers. The weights are not published, so there is nothing to download and run without a commercial agreement.

Does DeepSeek-OCR need a GPU?

In practice, yes. DeepSeek-OCR is a vision-language model of roughly 3 billion parameters, released under the MIT License in October 2025, and its model card lists GPU inference with CUDA. The project reports PDF throughput of about 2,500 tokens per second on a single A100 40 GB through vLLM, which supports the model upstream. DeepSeek-OCR-2, published in January 2026 with Apache 2.0 weights, is a similar size. Sites without GPUs are better served by Tesseract or PaddleOCR.

What is intelligent document processing on premise?

Intelligent document processing, or IDP, is the pipeline that turns incoming documents into checked data: classify each file, read the text, recover layout and tables, extract fields, validate them against business rules, send doubtful cases to a person and pass the result to downstream systems. On premise, every one of those steps runs on hardware you control, including the models, the review queue and the logs, so neither the page images nor the extracted values leave your environment.

Which document AI tools can read handwriting?

Check each project's own claims and then test on your forms. Tesseract's FAQ says handwriting will not work very well because the engine is designed for printed text. olmOCR lists handwriting among its supported content, and PaddleOCR describes its PaddleOCR-VL model as suited to handwritten text and historical documents. Among commercial options, Azure's Read and Layout models support handwriting in 12 languages in version 4.0, and Mistral listed handwriting among the strengths of its OCR 3 model.

Filed under
on-premises AIdocument automationopen-source AIAI procurementdata sovereigntyprivate RAG
Private RAG & Search

Evaluate your knowledge stack

Find out how a private RAG and retrieval layer would perform on your data — accuracy, latency, governance, and what to fix before you scale.

Or start free — no credit card →

Keep reading