AI Infrastructure

Mistral On-Premise: Open Models, Self-Deployment and Licensing (2026)

Mistral AI publishes some of its strongest models under Apache 2.0 and keeps others behind a revenue-capped or commercial licence. This guide sorts the October 2026 lineup by licence, explains Mistral's own self-deployment offers and data terms, sizes the GPUs each open model needs, and gives the vLLM, llama.cpp and Ollama steps to serve them.

Mistral on premise means running Mistral AI's models on infrastructure you control. Mistral Large 3, Mistral Small 4 and the Ministral 3 family are Apache 2.0 open weights you can self-host freely. Mistral Medium 3.5 needs a commercial licence above US$20 million monthly revenue, and proprietary models such as OCR need an agreement with Mistral.

Mistral is the European vendor that most often comes up when a regulated organisation asks for open models it can run itself. The lineup has grown into several licence regimes, though, and Mistral now sells its own self-hosted products next to the open weights. Knowing which model sits under which terms is the first decision, and the GPU plan follows from it.

This guide covers the current lineup, the licences, Mistral’s deployment offers and data terms, sizing and serving. Facts were checked against Mistral’s documentation, licence files, legal pages and Hugging Face repositories (verified October 2026). Companion guides cover Gemma 4 requirements and gpt-oss setup.

Mistral models you can self-host: quick picks

NeedModelLicenceSize and footprintStarting hardware
A capable model on one workstationMinistral 3 14BApache 2.014B dense; FP8 15.7 GB, Q4_K_M 8.2 GBOne 16 to 24 GB GPU
Edge and small devicesMinistral 3 3B or 8BApache 2.0FP8 4.7 GB or 10.4 GBLaptops, small GPUs
A team model with reasoning and imagesMistral Small 4Apache 2.0119B MoE, 6.5B active; FP8 120.9 GBTwo H200s, or one in NVFP4
A frontier-class open modelMistral Large 3Apache 2.0675B MoE, 41B active; FP8 681.5 GBOne 8-GPU H200 or B200 node
Mistral’s dense flagshipMistral Medium 3.5Modified MIT, revenue cap128B dense; FP8 133.6 GBFour 80 GB GPUs

Mistral’s open-weight lineup in October 2026

Mistral’s model overview marks each model as open or commercial. The open generalist models, all on Hugging Face under Apache 2.0:

  • Mistral Large 3, released on 2 December 2025. A granular mixture-of-experts model with 675B total and 41B active parameters, including a 2.5B vision encoder, and a 256K context. Mistral publishes FP8, NVFP4 and BF16 checkpoints plus an EAGLE draft model for speculative decoding.
  • Mistral Small 4, announced on 16 March 2026. 119B parameters with 6.5B active per token, 128 experts with 4 active, text and image input, and a 256K context. Mistral says it unifies the reasoning of Magistral, the multimodal skills of Pixtral and the agent work of Devstral in one model, with a per-request reasoning_effort of none or high.
  • Ministral 3 at 3B, 8B and 14B, each in base, instruct and reasoning versions, with image input and a 256K context. Mistral says the 14B fits in 24 GB of VRAM in FP8.

Specialist open models include Devstral Small 2 (24B) for software-engineering agents, Magistral Small 1.2 (24B), the Voxtral Small and Mini speech models, Voxtral Mini Transcribe Realtime, and the 3B Shieldstral 1.0 moderation model.

One operational detail favours self-hosting. Mistral retires hosted model versions on a schedule: its overview lists Devstral Small 2 as retired from the API on 31 March 2026, Devstral 2 on 31 July 2026, and the Magistral versions by 31 July 2026. The open weights stay downloadable, so a self-hosted deployment keeps running a version your tests approved until you choose to move.

Which Mistral models need a commercial licence

Five licence regimes run through the catalogue. We read each licence file in October 2026:

LicenceCheckpointsWhat it allows
Apache 2.0Large 3, Small 4, Ministral 3, Devstral Small 2, Magistral Small, Voxtral Small and Mini, Shieldstral 1.0, older Mistral 7B, Mixtral and NeMoCommercial use, modification and redistribution with the licence and notices kept
Modified MITMistral Medium 3.5, Devstral 2 (123B)Broad MIT rights, but none at all if your company’s global monthly revenue exceeded US$20 million in the preceding month; derivatives are covered too
Mistral Research LicenseMistral Large 2411, Pixtral Large 2411, Ministral 8B 2410Research purposes only; other uses need a licence from Mistral
Mistral AI Non-Production LicenseCodestral 22BTesting, research, personal or evaluation use in non-production environments
CC BY-NC 4.0Voxtral TTSNon-commercial use

Then there are models with no public weights at all: Codestral 25.08, Mistral Embed, Mistral Moderation 2, Voxtral Mini Transcribe 2 and the OCR line are proprietary in Mistral’s overview. The OCR models matter most for document workloads; our document AI comparison covers Mistral’s self-hosted container offer for them.

The Medium 3.5 clause is the one that catches large organisations. Above the threshold you can ask Mistral for a commercial licence, which it may grant at its sole discretion, or use the model through Mistral’s hosted services. Check revenue at group level, because the clause counts your company or your employer. Our note on licence clauses that create work covers the wider review.

Mistral’s own self-deployment offers

Mistral’s deployment documentation draws the line plainly: open-weight Apache 2.0 models “can be deployed on compatible hardware”, while commercial models come through cloud providers or Mistral Compute. For running open models yourself, it lists vLLM, TensorRT-LLM, TGI, SkyPilot, Cerebrium and Cloudflare Workers AI, with configurations from a single GPU to multi-node clusters.

Beyond the weights, Mistral sells three ways to run its software on your side of the line (verified October 2026):

  • Vibe, formerly Le Chat. Mistral’s product page says enterprise customers can deploy Vibe on-premises, in a private cloud or on Mistral Cloud with full data residency. Custom deployments sit in the Enterprise tier, priced on request.
  • Mistral Studio. The production platform offers hybrid, dedicated and self-hosted deployment; for self-hosting, Mistral’s wording is that nothing leaves your perimeter.
  • Partner-served deployments. Mistral’s partner-served deployment terms, updated 28 May 2026, cover Mistral products run through a cloud provider or reseller, including on your own infrastructure. The cloud provider becomes your sole point of contact for support, and you may not market the products as a standalone offering without Mistral’s written consent.

These routes matter for the proprietary and revenue-capped models. For Apache 2.0 weights you need no agreement with Mistral at all, though a support contract may still be worth having.

The EU-sovereignty angle, stated factually

Mistral is often bought as the European option. The verifiable facts behind that:

  • Corporate home. Mistral’s commercial terms identify it as a French company with registered offices in Paris, and apply French law with Paris courts for customers outside the Americas and Asia-Pacific (terms effective 25 September 2026).
  • Hosted data location. Mistral’s help centre says data is hosted in the EU by default, in the US only if you choose the US endpoint, and that some features can transfer data outside the EU under GDPR Article 46 safeguards. Enterprise customers can switch off some of those features.
  • Regional inference. Mistral made Regional Endpoints generally available in August 2026, letting customers keep inference in Europe or the US. The same announcement added third-party open models, starting with Z.ai’s GLM-5.2, to Mistral’s platform.
  • European compute. Mistral Compute, announced in June 2025, is Mistral’s own AI infrastructure in Europe, pitched as an alternative to US or China-based cloud and AI providers.

On premises, those points shrink to two. A self-hosted Apache 2.0 checkpoint sends nothing to Mistral, so data location is a property of your own network. And the licence has no domicile clause, unlike Meta’s Llama 4 multimodal grant, which excludes EU-headquartered companies as our Llama licence review explains. The vendor’s nationality still matters for support contracts, the commercial licences above and any hosted services you add.

GPU sizing for Mistral models

Two attention designs split the lineup. Large 3 and Small 4 use multi-head latent attention, caching a compressed latent of 576 and 320 values per layer. Medium 3.5 and Ministral 3 use grouped-query attention with 8 KV heads of dimension 128 in every layer, so their caches grow much faster. Layer counts come from each repository’s params.json; weights are the published checkpoint sizes. KV figures assume a 16-bit cache, following our GPU sizing method:

ModelWeights as publishedKV per tokenOne 32K sessionOne 128K session
Ministral 3 3B4.7 GB FP8104 KiB≈3.5 GB≈14.0 GB
Ministral 3 8B10.4 GB FP8136 KiB≈4.6 GB≈18.3 GB
Ministral 3 14B15.7 GB FP8; 8.2 GB Q4_K_M160 KiB≈5.4 GB≈21.5 GB
Mistral Small 4120.9 GB FP8; 70.8 GB NVFP422.5 KiB≈0.75 GB≈3.0 GB
Mistral Medium 3.5133.6 GB FP8352 KiB≈11.8 GB≈47.2 GB
Mistral Large 3681.5 GB FP8; 403.1 GB NVFP468.6 KiB≈2.3 GB≈9.2 GB

Dividing by 0.9 to keep a tenth of memory for activations and the runtime, with ten sessions standing for 50 named users at 20% peak concurrency:

ScenarioNeededResult
Ministral 3 14B, Q4_K_M, one 16K session≈12.1 GBFits a 16 GB GPU
Ministral 3 14B, FP8, one 16K session≈20.5 GBFits a 24 GB GPU; a 32K session is tight
Mistral Small 4, FP8, ten 32K sessions≈142.7 GBTwo 141 GB H200s at half capacity
Mistral Small 4, NVFP4, ten 32K sessions≈87.1 GBOne H200
Mistral Medium 3.5, FP8, ten 32K sessions≈279.7 GBFour 80 GB H100s; two H200s only with an FP8 cache
Mistral Large 3, FP8, twenty 32K sessions≈808 GBOne 8× H200 node at 71%
Mistral Large 3, NVFP4, twenty 32K sessions≈499 GBOne 8× H100 node at 77%

Medium 3.5 stores about fifteen times as much cache per token as Small 4, which is why ten concurrent sessions nearly double its footprint. Mistral’s own minimum for Small 4 is “4x NVIDIA HGX H100, 2x NVIDIA HGX H200, or 1x NVIDIA DGX B200”, and it says Medium 3.5 runs self-hosted “on as few as four GPUs”. Before planning around NVFP4, confirm that your GPU generation and engine build support the format.

Serving Mistral models with vLLM, llama.cpp and Ollama

Mistral’s current model cards point to vLLM for production serving. The steps below follow those cards; our engine comparison covers when a single-user runtime is enough.

  1. Install vLLM. For Small 4, uv pip install -U vllm pulls in mistral_common 1.11.0 or later. Medium 3.5 needs a vLLM nightly build with Transformers 5.4.0 or later. The current repositories are ungated.
  2. Serve Small 4 with the card’s command: vllm serve mistralai/Mistral-Small-4-119B-2603 --max-model-len 262144 --tensor-parallel-size 2 --attention-backend FLASH_ATTN_MLA --tool-call-parser mistral --enable-auto-tool-choice --reasoning-parser mistral. Lower --max-model-len to the context you actually use.
  3. Serve Large 3 across a full node: vllm serve mistralai/Mistral-Large-3-675B-Instruct-2512 --max-model-len 262144 --tensor-parallel-size 8 --tokenizer_mode mistral --config_format mistral --load_format mistral --enable-auto-tool-choice --tool-call-parser mistral.
  4. Serve Medium 3.5 with --tensor-parallel-size 8 and the same mistral tool and reasoning parsers, and add its EAGLE draft model for faster decoding. Mistral fixed a long-context error in the Transformers config after release; GGUF files built from the old config are affected, so rebuild or re-download them.
  5. Run Ministral 3 on one machine. Mistral publishes official GGUF files at Q4_K_M, Q5_K_M, Q8_0 and BF16 for llama.cpp, and Ollama carries the family: ollama run ministral-3:14b pulls a 9.1 GB build. Ollama lists Medium 3.5 locally at 80 GB in Q4_K_M, but its Mistral Large 3 tag is cloud-only.
  6. Set reasoning and sampling per request. Small 4 and Medium 3.5 take reasoning_effort of none or high, with temperature 0.7 for high. For Large 3 and Ministral 3, the cards recommend a temperature below 0.1 in production.
  7. Close the network. Serve from a segment without outbound access, set HF_HUB_OFFLINE=1 once weights are cached (Hugging Face docs) and opt out of vLLM usage statistics with VLLM_NO_USAGE_STATS=1 (vLLM docs).

How VDF AI fits

Mistral is one of the open-weight families VDF AI Chat names as common choices for in-perimeter serving, and VDF AI Agents can use any OpenAI-compatible endpoint, so a vLLM server running Mistral Small 4 or Ministral 3 can back individual agents under their own policies.

VDF AI Router applies policy before its learned routing: pinned models per workload, organisation-wide allow and deny lists, and regulated domains that only consider approved models. A deny rule can keep a revenue-capped checkpoint such as Medium 3.5 out of every workflow until the licence is settled, and air-gap mode restricts routing to local models. The Model Evaluation Suite then compares candidate Mistral models on your stored test cases inside your deployment and keeps the scored results for the approval decision.

Sources

Frequently asked questions

Can you run Mistral AI models on premises?

Yes. Mistral publishes open weights for Mistral Large 3, Mistral Small 4, the Ministral 3 family, Devstral Small 2 and several speech and safety models under Apache 2.0, and its documentation says these can be deployed on your own hardware with engines such as vLLM, TensorRT-LLM and TGI. Mistral Medium 3.5 also has public weights, but its licence grants nothing to a company whose monthly revenue tops 20 million dollars. Proprietary models such as OCR 4.1 need a commercial arrangement.

Which Mistral models are Apache 2.0?

As of October 2026, Mistral Large 3, Mistral Small 4, Ministral 3 at 3B, 8B and 14B, Devstral Small 2, Magistral Small, Voxtral Small and Mini, Voxtral Mini Transcribe Realtime and Shieldstral 1.0 carry Apache 2.0 on Hugging Face. Older Mistral 7B, Mixtral and Mistral NeMo checkpoints are Apache 2.0 too. Mistral Medium 3.5 and Devstral 2 use a Modified MIT licence, older Large 2 and Pixtral Large checkpoints use the research-only Mistral Research License, and Codestral 22B is non-production only.

Do I need a commercial licence for Mistral Medium 3.5?

Most enterprises do. The Modified MIT licence grants broad rights, but says you may not exercise any of them if the global consolidated monthly revenue of your company, or your employer, exceeded 20 million dollars in the preceding month. The restriction covers derivatives and fine-tunes too. Above that line you can request a commercial licence from Mistral, which it may grant at its discretion, or use the model through Mistral's hosted services. Smaller companies can self-host it under the licence as published.

How many GPUs does Mistral Large 3 need?

One eight-GPU node. Mistral recommends its FP8 checkpoint, about 682 GB, on a single node of B200s or H200s, and its NVFP4 checkpoint, about 403 GB, on a single node of H100s or A100s. By our arithmetic, eight H200s hold the FP8 weights plus twenty 32K-token sessions at about 71 percent of memory. Its latent attention keeps the cache small, roughly 2.3 GB per 32K-token session at 16-bit, so the weights decide the hardware.

Is Mistral a sovereign European alternative?

Mistral AI is a French company with its registered office in Paris, and its commercial terms apply French law for customers outside the Americas and Asia-Pacific. Its hosted services keep data in the EU by default, although some features can transfer data outside the EU. When you self-host Apache 2.0 weights, sovereignty comes from your own infrastructure rather than the vendor's passport: no data reaches Mistral at all, and the licence imposes no domicile restrictions.

Filed under
Mistralopen-weight modelslocal LLMdata sovereigntysovereign AImixture of expertson-premises AI
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading