Mistral on premise means running Mistral AI's models on infrastructure you control. Mistral Large 3, Mistral Small 4 and the Ministral 3 family are Apache 2.0 open weights you can self-host freely. Mistral Medium 3.5 needs a commercial licence above US$20 million monthly revenue, and proprietary models such as OCR need an agreement with Mistral.
Mistral is the European vendor that most often comes up when a regulated organisation asks for open models it can run itself. The lineup has grown into several licence regimes, though, and Mistral now sells its own self-hosted products next to the open weights. Knowing which model sits under which terms is the first decision, and the GPU plan follows from it.
This guide covers the current lineup, the licences, Mistral’s deployment offers and data terms, sizing and serving. Facts were checked against Mistral’s documentation, licence files, legal pages and Hugging Face repositories (verified October 2026). Companion guides cover Gemma 4 requirements and gpt-oss setup.
Mistral models you can self-host: quick picks
| Need | Model | Licence | Size and footprint | Starting hardware |
|---|---|---|---|---|
| A capable model on one workstation | Ministral 3 14B | Apache 2.0 | 14B dense; FP8 15.7 GB, Q4_K_M 8.2 GB | One 16 to 24 GB GPU |
| Edge and small devices | Ministral 3 3B or 8B | Apache 2.0 | FP8 4.7 GB or 10.4 GB | Laptops, small GPUs |
| A team model with reasoning and images | Mistral Small 4 | Apache 2.0 | 119B MoE, 6.5B active; FP8 120.9 GB | Two H200s, or one in NVFP4 |
| A frontier-class open model | Mistral Large 3 | Apache 2.0 | 675B MoE, 41B active; FP8 681.5 GB | One 8-GPU H200 or B200 node |
| Mistral’s dense flagship | Mistral Medium 3.5 | Modified MIT, revenue cap | 128B dense; FP8 133.6 GB | Four 80 GB GPUs |
Mistral’s open-weight lineup in October 2026
Mistral’s model overview marks each model as open or commercial. The open generalist models, all on Hugging Face under Apache 2.0:
- Mistral Large 3, released on 2 December 2025. A granular mixture-of-experts model with 675B total and 41B active parameters, including a 2.5B vision encoder, and a 256K context. Mistral publishes FP8, NVFP4 and BF16 checkpoints plus an EAGLE draft model for speculative decoding.
- Mistral Small 4, announced on 16 March 2026. 119B parameters with 6.5B active per token, 128 experts with 4 active, text and image input, and a 256K context. Mistral says it unifies the reasoning of Magistral, the multimodal skills of Pixtral and the agent work of Devstral in one model, with a per-request
reasoning_effortofnoneorhigh. - Ministral 3 at 3B, 8B and 14B, each in base, instruct and reasoning versions, with image input and a 256K context. Mistral says the 14B fits in 24 GB of VRAM in FP8.
Specialist open models include Devstral Small 2 (24B) for software-engineering agents, Magistral Small 1.2 (24B), the Voxtral Small and Mini speech models, Voxtral Mini Transcribe Realtime, and the 3B Shieldstral 1.0 moderation model.
One operational detail favours self-hosting. Mistral retires hosted model versions on a schedule: its overview lists Devstral Small 2 as retired from the API on 31 March 2026, Devstral 2 on 31 July 2026, and the Magistral versions by 31 July 2026. The open weights stay downloadable, so a self-hosted deployment keeps running a version your tests approved until you choose to move.
Which Mistral models need a commercial licence
Five licence regimes run through the catalogue. We read each licence file in October 2026:
| Licence | Checkpoints | What it allows |
|---|---|---|
| Apache 2.0 | Large 3, Small 4, Ministral 3, Devstral Small 2, Magistral Small, Voxtral Small and Mini, Shieldstral 1.0, older Mistral 7B, Mixtral and NeMo | Commercial use, modification and redistribution with the licence and notices kept |
| Modified MIT | Mistral Medium 3.5, Devstral 2 (123B) | Broad MIT rights, but none at all if your company’s global monthly revenue exceeded US$20 million in the preceding month; derivatives are covered too |
| Mistral Research License | Mistral Large 2411, Pixtral Large 2411, Ministral 8B 2410 | Research purposes only; other uses need a licence from Mistral |
| Mistral AI Non-Production License | Codestral 22B | Testing, research, personal or evaluation use in non-production environments |
| CC BY-NC 4.0 | Voxtral TTS | Non-commercial use |
Then there are models with no public weights at all: Codestral 25.08, Mistral Embed, Mistral Moderation 2, Voxtral Mini Transcribe 2 and the OCR line are proprietary in Mistral’s overview. The OCR models matter most for document workloads; our document AI comparison covers Mistral’s self-hosted container offer for them.
The Medium 3.5 clause is the one that catches large organisations. Above the threshold you can ask Mistral for a commercial licence, which it may grant at its sole discretion, or use the model through Mistral’s hosted services. Check revenue at group level, because the clause counts your company or your employer. Our note on licence clauses that create work covers the wider review.
Mistral’s own self-deployment offers
Mistral’s deployment documentation draws the line plainly: open-weight Apache 2.0 models “can be deployed on compatible hardware”, while commercial models come through cloud providers or Mistral Compute. For running open models yourself, it lists vLLM, TensorRT-LLM, TGI, SkyPilot, Cerebrium and Cloudflare Workers AI, with configurations from a single GPU to multi-node clusters.
Beyond the weights, Mistral sells three ways to run its software on your side of the line (verified October 2026):
- Vibe, formerly Le Chat. Mistral’s product page says enterprise customers can deploy Vibe on-premises, in a private cloud or on Mistral Cloud with full data residency. Custom deployments sit in the Enterprise tier, priced on request.
- Mistral Studio. The production platform offers hybrid, dedicated and self-hosted deployment; for self-hosting, Mistral’s wording is that nothing leaves your perimeter.
- Partner-served deployments. Mistral’s partner-served deployment terms, updated 28 May 2026, cover Mistral products run through a cloud provider or reseller, including on your own infrastructure. The cloud provider becomes your sole point of contact for support, and you may not market the products as a standalone offering without Mistral’s written consent.
These routes matter for the proprietary and revenue-capped models. For Apache 2.0 weights you need no agreement with Mistral at all, though a support contract may still be worth having.
The EU-sovereignty angle, stated factually
Mistral is often bought as the European option. The verifiable facts behind that:
- Corporate home. Mistral’s commercial terms identify it as a French company with registered offices in Paris, and apply French law with Paris courts for customers outside the Americas and Asia-Pacific (terms effective 25 September 2026).
- Hosted data location. Mistral’s help centre says data is hosted in the EU by default, in the US only if you choose the US endpoint, and that some features can transfer data outside the EU under GDPR Article 46 safeguards. Enterprise customers can switch off some of those features.
- Regional inference. Mistral made Regional Endpoints generally available in August 2026, letting customers keep inference in Europe or the US. The same announcement added third-party open models, starting with Z.ai’s GLM-5.2, to Mistral’s platform.
- European compute. Mistral Compute, announced in June 2025, is Mistral’s own AI infrastructure in Europe, pitched as an alternative to US or China-based cloud and AI providers.
On premises, those points shrink to two. A self-hosted Apache 2.0 checkpoint sends nothing to Mistral, so data location is a property of your own network. And the licence has no domicile clause, unlike Meta’s Llama 4 multimodal grant, which excludes EU-headquartered companies as our Llama licence review explains. The vendor’s nationality still matters for support contracts, the commercial licences above and any hosted services you add.
GPU sizing for Mistral models
Two attention designs split the lineup. Large 3 and Small 4 use multi-head latent attention, caching a compressed latent of 576 and 320 values per layer. Medium 3.5 and Ministral 3 use grouped-query attention with 8 KV heads of dimension 128 in every layer, so their caches grow much faster. Layer counts come from each repository’s params.json; weights are the published checkpoint sizes. KV figures assume a 16-bit cache, following our GPU sizing method:
| Model | Weights as published | KV per token | One 32K session | One 128K session |
|---|---|---|---|---|
| Ministral 3 3B | 4.7 GB FP8 | 104 KiB | ≈3.5 GB | ≈14.0 GB |
| Ministral 3 8B | 10.4 GB FP8 | 136 KiB | ≈4.6 GB | ≈18.3 GB |
| Ministral 3 14B | 15.7 GB FP8; 8.2 GB Q4_K_M | 160 KiB | ≈5.4 GB | ≈21.5 GB |
| Mistral Small 4 | 120.9 GB FP8; 70.8 GB NVFP4 | 22.5 KiB | ≈0.75 GB | ≈3.0 GB |
| Mistral Medium 3.5 | 133.6 GB FP8 | 352 KiB | ≈11.8 GB | ≈47.2 GB |
| Mistral Large 3 | 681.5 GB FP8; 403.1 GB NVFP4 | 68.6 KiB | ≈2.3 GB | ≈9.2 GB |
Dividing by 0.9 to keep a tenth of memory for activations and the runtime, with ten sessions standing for 50 named users at 20% peak concurrency:
| Scenario | Needed | Result |
|---|---|---|
| Ministral 3 14B, Q4_K_M, one 16K session | ≈12.1 GB | Fits a 16 GB GPU |
| Ministral 3 14B, FP8, one 16K session | ≈20.5 GB | Fits a 24 GB GPU; a 32K session is tight |
| Mistral Small 4, FP8, ten 32K sessions | ≈142.7 GB | Two 141 GB H200s at half capacity |
| Mistral Small 4, NVFP4, ten 32K sessions | ≈87.1 GB | One H200 |
| Mistral Medium 3.5, FP8, ten 32K sessions | ≈279.7 GB | Four 80 GB H100s; two H200s only with an FP8 cache |
| Mistral Large 3, FP8, twenty 32K sessions | ≈808 GB | One 8× H200 node at 71% |
| Mistral Large 3, NVFP4, twenty 32K sessions | ≈499 GB | One 8× H100 node at 77% |
Medium 3.5 stores about fifteen times as much cache per token as Small 4, which is why ten concurrent sessions nearly double its footprint. Mistral’s own minimum for Small 4 is “4x NVIDIA HGX H100, 2x NVIDIA HGX H200, or 1x NVIDIA DGX B200”, and it says Medium 3.5 runs self-hosted “on as few as four GPUs”. Before planning around NVFP4, confirm that your GPU generation and engine build support the format.
Serving Mistral models with vLLM, llama.cpp and Ollama
Mistral’s current model cards point to vLLM for production serving. The steps below follow those cards; our engine comparison covers when a single-user runtime is enough.
- Install vLLM. For Small 4,
uv pip install -U vllmpulls inmistral_common1.11.0 or later. Medium 3.5 needs a vLLM nightly build with Transformers 5.4.0 or later. The current repositories are ungated. - Serve Small 4 with the card’s command:
vllm serve mistralai/Mistral-Small-4-119B-2603 --max-model-len 262144 --tensor-parallel-size 2 --attention-backend FLASH_ATTN_MLA --tool-call-parser mistral --enable-auto-tool-choice --reasoning-parser mistral. Lower--max-model-lento the context you actually use. - Serve Large 3 across a full node:
vllm serve mistralai/Mistral-Large-3-675B-Instruct-2512 --max-model-len 262144 --tensor-parallel-size 8 --tokenizer_mode mistral --config_format mistral --load_format mistral --enable-auto-tool-choice --tool-call-parser mistral. - Serve Medium 3.5 with
--tensor-parallel-size 8and the samemistraltool and reasoning parsers, and add its EAGLE draft model for faster decoding. Mistral fixed a long-context error in the Transformers config after release; GGUF files built from the old config are affected, so rebuild or re-download them. - Run Ministral 3 on one machine. Mistral publishes official GGUF files at Q4_K_M, Q5_K_M, Q8_0 and BF16 for llama.cpp, and Ollama carries the family:
ollama run ministral-3:14bpulls a 9.1 GB build. Ollama lists Medium 3.5 locally at 80 GB in Q4_K_M, but its Mistral Large 3 tag is cloud-only. - Set reasoning and sampling per request. Small 4 and Medium 3.5 take
reasoning_effortofnoneorhigh, with temperature 0.7 forhigh. For Large 3 and Ministral 3, the cards recommend a temperature below 0.1 in production. - Close the network. Serve from a segment without outbound access, set
HF_HUB_OFFLINE=1once weights are cached (Hugging Face docs) and opt out of vLLM usage statistics withVLLM_NO_USAGE_STATS=1(vLLM docs).
How VDF AI fits
Mistral is one of the open-weight families VDF AI Chat names as common choices for in-perimeter serving, and VDF AI Agents can use any OpenAI-compatible endpoint, so a vLLM server running Mistral Small 4 or Ministral 3 can back individual agents under their own policies.
VDF AI Router applies policy before its learned routing: pinned models per workload, organisation-wide allow and deny lists, and regulated domains that only consider approved models. A deny rule can keep a revenue-capped checkpoint such as Medium 3.5 out of every workflow until the licence is settled, and air-gap mode restricts routing to local models. The Model Evaluation Suite then compares candidate Mistral models on your stored test cases inside your deployment and keeps the scored results for the approval decision.
Sources
- Mistral models overview and deployment documentation
- Mistral 3 announcement and Mistral Small 4 announcement
- Mistral Medium 3.5 docs model card and Vibe and Medium 3.5 announcement
- Hugging Face model cards: Mistral Large 3, Mistral Small 4, Mistral Medium 3.5 and Ministral 3 14B
- Ministral 3 14B GGUF files
- Licences: Medium 3.5 Modified MIT, Devstral 2 Modified MIT, Mistral Research License and Mistral AI Non-Production License
- Mistral Vibe and Mistral Studio
- Partner-served deployment terms and commercial terms of service
- Mistral help centre: where data is stored
- Mistral: regional inference and new compute and Mistral Compute
- Ollama: ministral-3 tags, mistral-medium-3.5 tags and mistral-large-3 tags
- Hugging Face Hub environment variables and vLLM usage statistics
- NVIDIA H100 and H200