On-premise document AI is OCR, layout analysis and field extraction running on infrastructure you control. Of the cloud services compared here, Azure Document Intelligence is the one you can run in your own containers, though a fully disconnected setup needs Microsoft's approval and a commitment plan. Google and SAP run Document AI as cloud services, Mistral self-hosts OCR for selected enterprises, and open models cover the rest.
What on-premise document AI has to do
Intelligent document processing (IDP) is a pipeline. On premise, each stage needs a component you run yourself:
- Intake and classification. Decide what each file is (invoice, ID, contract, claim form) and whether it already carries a usable text layer.
- Text recognition. OCR for scans and photographs, direct text extraction for digitally produced PDFs.
- Layout and tables. Reading order, headings, table cells, checkboxes and key-value pairs.
- Field extraction. Named values for each document class, from templates, trained extractors or a language model.
- Validation and review. Business rules, confidence thresholds and a human queue for doubtful fields.
- Hand-off. Structured output into an ERP, a case system or a search index, linked back to the page image.
Cloud IDP services bundle most of these stages behind one API. Self-hosting means choosing a component per stage: more integration work, in exchange for control over every model, log and retention rule. The retrieval side of the same problem, including chunking, provenance and confidence routing for private RAG, is covered in our scanned-documents deep dive. If invoices are the first document class, our post on invoice exception agents picks up after extraction.
Cloud document AI services with an on-premise route
| Service | On-premise option (verified October 2026) | Conditions |
|---|---|---|
| Azure Document Intelligence | Docker containers, connected or disconnected | Connected mode meters usage through Azure; disconnected mode needs approval, a commitment plan and an expiring license file |
| Google Document AI | None in the Document AI documentation | Processors run in Google Cloud locations; Google Distributed Cloud air-gapped offers a separate OCR API |
| SAP Document AI | None found | Managed service on SAP BTP that reaches on-premise SAP systems through SAP Cloud Connector |
| Mistral OCR 4 and 4.1 | Self-hosted single container for enterprise customers | Proprietary model; API at $4 per 1,000 pages, $2 with the Batch API (list, verified October 2026) |
Azure Document Intelligence containers
Microsoft publishes Azure Document Intelligence in Foundry Tools as Docker images in the Microsoft Container Registry. Which models you can run depends on the API version:
| API version | Models available as containers |
|---|---|
| v4.0 (2024-11-30) | Read, Layout |
| v3.1 (2023-07-31) | Read, Layout, ID Document, Receipt, Invoice |
| v3.0 (2022-08-31) | Read, Layout, General Document, Business Card, Custom (template) |
The newest containers therefore cover OCR and layout only. Prebuilt invoice, receipt and ID extraction means running v3.1 images, custom extraction means v3.0 template models, and Microsoft’s documentation already marks v3.0 as retiring. Several containers also depend on another one: Invoice needs Layout, Receipt needs Read or Layout, and ID Document and Business Card need Read.
Microsoft’s sizing table for disconnected use asks for 8 cores per container, with minimum memory from 8 GB for ID Document to 16 GB for Layout, Invoice and Custom Template, and 24 GB recommended across the board. Connected containers report usage to Azure for billing, and Microsoft states that Azure containers do not send customer data such as the analysed image or text. Disconnected use starts with a request form and a commitment-tier resource. The container then downloads a license file while still connected, and that file carries an expiry date. On handwriting, the v4.0 Read and Layout models list 12 languages: English, Chinese Simplified, French, German, Italian, Thai, Japanese, Korean, Portuguese, Spanish, Russian and Arabic.
Google Document AI and Google Distributed Cloud
Google describes Document AI as a platform that turns unstructured document data into structured data. Processors are created in the us or eu multi-regions or in one of seven single regions with limited support, and the Document AI documentation lists no container or on-premises edition.
For disconnected sites, Google Distributed Cloud air-gapped brings Vertex AI into on-premises environments. Its three pre-trained APIs are OCR, Speech-to-Text and Translation. The OCR service detects printed and handwritten text in JPG, PNG, PDF and TIFF files, but Document AI’s processors are not on that list, so treat it as a text-recognition API inside Google’s air-gapped platform.
SAP Document AI
SAP’s reference architecture describes Document AI as a managed service on SAP Business Technology Platform that combines OCR, pretrained and customer-trained extraction models and LLM-based reasoning. SAP Cloud Connector gives it a path into on-premise systems such as SAP ECC and SAP S/4HANA, so extracted values can land in an on-premise ERP while the documents themselves are processed on SAP BTP.
Mistral OCR
Mistral announced OCR 4 on 23 June 2026, and OCR 4.1 became generally available on 16 July 2026. Both are proprietary models in Mistral’s Premier tier, priced as shown in the table above. Mistral says OCR 4 is compact enough to deploy in a single container on your own infrastructure and offers that self-managed deployment to enterprise customers. OCR 4 returns bounding boxes, block types such as tables and signatures, and per-page and per-word confidence scores, and Mistral claims support for 170 languages across 10 language groups.
Open OCR and document parsing models compared
These projects run wherever you have the hardware. Licences, features and requirements come from each project’s repository or model card (verified October 2026). “Not stated” means the project makes no claim either way, so test it yourself.
| Project | Approach | Licence | Layout and tables | Handwriting | Languages | Hardware |
|---|---|---|---|---|---|---|
| Tesseract | LSTM engine focused on line recognition | Apache 2.0 | Text with positions; hOCR, ALTO, PAGE, TSV and searchable PDF output | FAQ: will not work very well; designed for printed text | 100+ out of the box | CPU |
| PaddleOCR | PP-OCRv5 and PP-OCRv6 pipelines, PP-StructureV3 parsing, the 0.9B PaddleOCR-VL model | Apache 2.0 | PP-StructureV3 returns table cell and text coordinates as Markdown or JSON | PaddleOCR-VL described as suited to handwritten and historical documents | PP-OCRv5: 109; PP-OCRv6: 50 in one model; PaddleOCR-VL-1.5: 111 | CPU, GPU, XPU and NPU; ONNX Runtime, OpenVINO and TensorRT |
| docTR | Two-stage text detection and recognition on PyTorch | Apache 2.0 | Optional layout detection (titles, text, tables, page headers and footers) and a key information extraction predictor | Not stated | Character vocabularies defined for dozens of languages | PyTorch |
| Docling | Converter with layout and table-structure models, plus an optional VLM pipeline | MIT; Granite-Docling-258M weights Apache 2.0 | Page layout, reading order, table structure, code and formulas | Not stated | Granite-Docling: English, with experimental Japanese, Arabic and Chinese | Documented for local and air-gapped execution |
| Marker | Converts PDF, images and Office files to Markdown, JSON, HTML or chunks | Code Apache 2.0; weights under a modified OpenRAIL-M licence | Table converter; optional LLM mode merges tables across pages | Not stated | Claims all languages | GPU, CPU or Apple MPS |
| olmOCR | Toolkit around a 7B VLM fine-tuned from Qwen2.5-VL | Apache 2.0 | Tables, equations and multi-column reading order; strips headers and footers | Listed as supported | Not stated | NVIDIA GPU with at least 12 GB; 30 GB of free disk |
| DeepSeek-OCR and DeepSeek-OCR-2 | VLMs of about 3B parameters, steered by prompts | MIT (OCR); Apache 2.0 (OCR-2) | Prompt modes for document-to-Markdown with layout grounding and for parsing figures | Not stated | Not stated | NVIDIA GPU; supported in upstream vLLM |
A few notes the table cannot carry:
- Tesseract suits clean printed scans, and its maintainers say plainly that better results often depend on improving image quality first.
- PaddleOCR moves quickly. Release 3.7.0 shipped in June 2026, and PP-OCRv6 recognises its 50 languages with a single unified model, so multilingual pages need no model switching.
- Docling is the parser the scanned-documents post examines in depth; for this comparison, what matters is its MIT licence and its explicit support for air-gapped environments.
- olmOCR release notes for version 0.3.0 mention fixing hallucinations on blank documents, a reminder that VLM-based OCR fails differently from classic engines.
- DeepSeek-OCR was released on 20 October 2025 and added upstream vLLM support three days later; DeepSeek-OCR-2 followed in January 2026.
Licence and data-flow traps to check before a pilot
- Marker’s model weights are not free for everyone. The modified OpenRAIL-M licence covers research, personal use and startups under $5M in funding or revenue; larger commercial users need a licence from Datalab.
- Marker’s LLM mode calls Gemini by default. The optional —use_llm flag improves tables and forms, and out of the box it uses a Google Gemini model. On premise, point it at a local model served through Ollama or leave it off.
- olmOCR can talk to hosted endpoints. Its README lists external inference providers it has tested; make sure the server flag points at your own vLLM instance.
- Docling and Granite-Docling fetch models on first use. The Granite-Docling card notes that Docling downloads the model automatically, so mirror the weights before moving into an air-gapped network.
- Azure’s disconnected licence expires. Plan for the commitment term and fresh license files when you pull new container images, because Microsoft warns an old file may fail to start a new image.
- Mistral’s self-hosted OCR is a commercial agreement. Budget for negotiation time, since there are no published weights to evaluate first.
Hardware for document AI on premise
Document AI splits into three hardware tiers:
- CPU tier. Tesseract runs on CPUs, and PaddleOCR supports Intel CPUs with OpenVINO acceleration, which suits high-volume printed scans with no GPU budget.
- Container tier. Azure’s containers run on x64 CPU hosts: plan 8 cores and up to 24 GB of memory per container, plus any supporting container the model needs.
- GPU tier. Vision-language models need accelerators. olmOCR asks for an NVIDIA GPU with at least 12 GB and was tested on RTX 4090, L40S, A100 and H100. DeepSeek-OCR’s PDF throughput figure comes from an A100 40 GB, and its roughly 3 billion parameters ship in BF16.
For GPU sizing, apply the weights and KV-cache arithmetic from our GPU estimation guide to the VLM you pick. Backfilling an archive is a long batch job that can starve interactive users, so give it its own queue with GPU admission control.
Accuracy risks with vision-language OCR
Classic OCR engines tend to fail visibly, with garbled characters on a poor scan. Vision-language models can fail fluently, producing clean text that is not on the page, as olmOCR’s fix for blank documents shows. Three habits contain both kinds of error:
- Ask for confidence. Mistral OCR 4 returns per-page and per-word confidence scores. Where an open model returns none, use proxies such as agreement between two engines, and route disagreements to review.
- Keep the page image beside every value. Reviewers should see the source region next to the extracted field, and the stored record should carry document, page and region identifiers.
- Test on your worst documents. Faded photocopies, rotated scans, stamps, handwriting and mixed languages expose differences that clean benchmark pages hide.
A selection checklist for on-premise IDP
- Inventory document classes and volumes. Pages per day, peak backfill size and which classes carry personal or regulated data.
- Split digital from scanned at intake. Digitally produced files rarely need OCR at all.
- Set a processing boundary per class. On premise only, disconnected, or an approved cloud region.
- Read the licences for code and weights. Look for revenue thresholds, non-commercial clauses and commitment terms.
- Match language and handwriting coverage. Compare each class against the vendor’s or project’s own published lists, not marketing summaries.
- Demand tables as structure. Cell coordinates and spans, not flattened text.
- Require confidence or build a proxy. Then set thresholds per field, not per document.
- Design human review into the flow. Side-by-side page image, field value and an audit trail of corrections.
- Keep provenance. Document ID, version, page and region for every extracted value.
- Size hardware per tier. Separate CPU, container and GPU workloads, and isolate batch backfills.
- Plan updates. Mirror models, pin versions, re-run your evaluation set, and track licence expiry for vendor containers.
- Log every run. Which document, which model version, which reviewer and what changed.
How VDF AI fits
VDF AI handles document work as tools that agents call inside your environment. The document parser tool turns files of most formats into clean text and structure, the table extraction tool recovers headers, rows and cells from documents, PDFs and images, and the OCR tool reads text from scans and photos. With the Tesseract engine selected, OCR runs entirely inside your perimeter with multi-language support and audit logging; a vision-model engine is available for difficult images.
For analysts, the report analysis product accepts PDF, Word, Excel, CSV and PowerPoint files as well as JPG, PNG and TIFF images read with OCR, and answers questions about them in a chat interface.
Sources
- Microsoft Learn, install and run Document Intelligence containers
- Microsoft Learn, Document Intelligence containers in disconnected environments
- Microsoft Learn, containers in disconnected environments (Foundry Tools)
- Microsoft Learn, Read and Layout language support
- Google Cloud, Document AI overview
- Google Cloud, Document AI locations
- Google, OCR on GDC air-gapped
- Google, Vertex AI on GDC air-gapped
- SAP Architecture Center, SAP Document AI
- Mistral, OCR 4 announcement
- Mistral docs, OCR 4.1
- Mistral, OCR 3 announcement
- Tesseract repository
- Tesseract FAQ
- PaddleOCR repository
- docTR repository
- Docling repository
- Granite-Docling-258M model card
- Marker repository
- olmOCR repository
- DeepSeek-OCR repository
- DeepSeek-OCR model card
- DeepSeek-OCR-2 model card