Llama on premise means running Meta's open-weight Llama models on hardware you control. As of October 2026 the current releases are Llama 4 Scout and Maverick, both from April 2025, and Llama 3.3 70B. Meta's community licences allow commercial self-hosting but add a 700-million-user threshold, naming and attribution rules, an acceptable use policy and an EU exclusion for multimodal models.
Llama is often the first open-weight family an enterprise tries, because every serving engine and most tooling support it. Three things still need a decision: licence clauses that other families lack, a Llama 4 design that trades memory for speed, and a line-up that has not changed in eighteen months. Facts below were checked on Hugging Face, Meta’s developer site and the vLLM documentation on 2 October 2026.
The Llama family as of October 2026
| Model | Released | Parameters | Context | Input |
|---|---|---|---|---|
| Llama 4 Scout | 5 Apr 2025 | 109B total, 17B active, 16 experts | 10M tokens | Text and images |
| Llama 4 Maverick | 5 Apr 2025 | 400B total, 17B active, 128 experts | 1M tokens | Text and images |
| Llama 3.3 70B | 6 Dec 2024 | 70B dense | 128K tokens | Text |
| Llama 3.2 | 25 Sep 2024 | 1B and 3B text; 11B and 90B Vision | 128K tokens | Text; text and images for Vision |
| Llama 3.1 | 23 Jul 2024 | 8B, 70B and 405B dense | 128K tokens | Text |
Both Llama 4 models have a knowledge cutoff of August 2024 and list twelve supported languages, among them English, French, German, Italian, Portuguese and Spanish. Meta’s launch post says Scout was pre-trained and post-trained with a 256K context length, so the 10M-token figure relies on length generalisation. Test the context you actually need before you plan around it.
Llama 4 Behemoth was described as still training at launch, and no Behemoth checkpoint has appeared in Meta’s Hugging Face organisation (verified October 2026). In July 2025 Mark Zuckerberg wrote that Meta would need to be careful about what it chooses to open source (Meta). Meta’s newer models carry the Muse name: Muse Spark launched in April 2026 without open weights, offered to developers through Meta’s API, and Muse Glimmer, released on 10 August 2026, is a 30B open-weight model under the Apache 2.0 licence, aimed at local agents on a single consumer GPU. The same day, Zuckerberg wrote that Meta “will resume releasing some open source models soon” (essay).
For planning, treat Llama 4 and Llama 3.3 as the Llama-branded options, and Muse Glimmer as a separate family with different licence terms.
What the Llama licence asks of an enterprise
Each Llama version has its own community licence and acceptable use policy. We read the Llama 4 licence, the Llama 3.3 licence and the Llama 4 acceptable use policy on 2 October 2026:
| Clause | What it says | What it means on premises |
|---|---|---|
| Acceptable use | The acceptable use policy is incorporated by reference | Fold its prohibited uses into your own AI acceptable-use policy |
| 700 million users | If your products had more than 700 million monthly active users in the month before the version’s release date, you must request a licence from Meta, which it may grant at its sole discretion | Irrelevant for almost every enterprise, but it is counted across affiliates |
| Redistribution | Ship a copy of the agreement, display “Built with Llama” prominently, start the names of derived models with “Llama” and include Meta’s attribution notice | Applies when you distribute Llama or a product built on it; sharing a fine-tuned model with another legal entity is likely to count |
| EU and multimodal models | For the multimodal models in Llama 4, the Section 1(a) rights are not granted to individuals domiciled in the EU or companies with their principal place of business there | Excludes EU-headquartered companies from self-hosting Scout and Maverick under the standard grant |
The EU clause needs the most attention. Both Llama 4 models accept images, so both are multimodal, and the exemption covers end users of a product that incorporates them, not a company running the weights itself. Llama 3.2’s policy has the same clause for its 11B and 90B Vision models, while Llama 3.3 70B and the text-only 3.1 and 3.2 models are outside it. An EU-headquartered organisation that wants Llama on its own GPUs is therefore choosing between Llama 3.3 70B and older text models; counsel should confirm how the clause applies to your group structure.
How Llama 4 uses memory
Two design choices separate Llama 4 from Llama 3.3. The figures come from the config.json files, read from ungated mirrors because Meta’s own repositories require approved access.
Mixture of experts. Scout has 16 experts in every one of its 48 layers; Maverick has 128 experts in alternating layers, 24 of the 48. Each token is routed to one expert, so both models compute with about 17B active parameters while keeping all 109B or 400B in GPU memory. A mixture-of-experts model therefore generates text at roughly the cost of a much smaller dense model, while its weights cost the full amount.
Mixed attention. Both models use 40 query heads and 8 key-value heads of dimension 128. In three of every four layers, attention is chunked: a token attends only within its 8,192-token chunk, using rotary position embeddings. Every fourth layer, 12 of the 48, attends across the whole context without positional encoding. Because the chunked layers never look past the current chunk, an engine only has to keep that chunk’s keys and values for them, and the config’s cache_implementation: hybrid setting reflects this.
Llama 3.3 70B is a conventional dense model: 80 layers, 64 query heads and 8 key-value heads of dimension 128, with full attention in every layer, so its KV cache grows by the same amount for every token at any context length.
Sizing GPUs for Llama
Two numbers decide the fit. Weights take roughly parameters × bytes per parameter. The KV cache takes 2 × layers × KV heads × head dimension × bytes per value for every token of every concurrent sequence, as in NVIDIA’s derivation and our GPU sizing method. The table applies that to the published configs at a 16-bit KV cache:
| Model | Weights: BF16 / FP8 / 4-bit | KV cache per token | Per 8,192-token sequence | Per 131,072-token sequence |
|---|---|---|---|---|
| Llama 3.1 8B | ≈16 / ≈8 / ≈5 GB | 128 KiB | ≈1.1 GB | ≈17 GB |
| Llama 3.3 70B | ≈141 / ≈71 / ≈35–40 GB | 320 KiB | ≈2.7 GB | ≈43 GB |
| Llama 4 Scout | ≈217 / ≈109 / ≈55–67 GB | 192 KiB; 48 KiB per token past the first 8,192 with a hybrid cache | ≈1.6 GB | ≈7.7 GB hybrid, ≈26 GB if every layer is cached in full |
| Llama 4 Maverick | ≈803 / ≈417 / ≈200–245 GB | Same as Scout | ≈1.6 GB | Same as Scout |
The BF16 sizes for Llama 4 and the FP8 size for Maverick are the totals in the safetensors index files; Meta publishes no FP8 build of Scout, so its FP8 figure is arithmetic. The 4-bit ranges run from half a byte per parameter up to the size of Ollama’s 4-bit builds, 67 GB for Scout and 245 GB for Maverick. GPU capacities are NVIDIA’s: 80 GB for an H100 SXM and 141 GB for an H200.
Worked example. Suppose 100 to 200 named users, which at 10–20% concurrency gives about 20 sequences at once, each with a 16,384-token context. Divide the total by 0.9 to leave a tenth of GPU memory for activations and the runtime.
- Llama 3.3 70B at FP8: 71 GB of weights plus 20 × 5.4 GB of KV cache is about 178 GB, or 198 GB with headroom. That is two H200s or four H100s. An FP8 KV cache brings it down to about 138 GB: one H200 at its limit, or two H100s.
- Llama 4 Scout at BF16, hybrid cache: each sequence holds about 0.8 GB for the global layers and 1.2 GB for one chunk, so 20 sequences need about 40 GB. With 217 GB of weights the total is about 286 GB with headroom, which means four H100s.
- Llama 4 Scout at 4-bit: around 58 GB of weights plus the same 40 GB of KV cache comes to about 109 GB with headroom: one H200, or two H100s.
Scout needs more memory to load but much less per long conversation, and it computes with 17B parameters per token against 70B. Llama 3.3 is cheaper to hold and simpler to serve. For the server around the GPUs, see the on-premise AI server guide.
Serving Llama with vLLM, llama.cpp and Ollama
vLLM lists Llama4ForConditionalGeneration among its supported models, with text and image input (supported models). Its Llama 4 launch notes required vLLM 0.8.3 or later and reported these limits:
- On eight H100s, Scout with a 1M-token context and Maverick with about 430K.
- On eight H200s, Scout up to 3.6M tokens and Maverick up to 1M.
- The commands used
--tensor-parallel-size 8and--max-model-len, served Maverick from its FP8 checkpoint, and setattn_temperature_tuningthrough--override-generation-configfor Scout’s long-context run. --kv-cache-dtype fp8can roughly double the usable context, at a quality cost you should measure.
One retirement matters for older deployments: vLLM’s model registry records that the Llama 3.2 Vision architecture “was supported in vLLM until v0.10.2”. A team still running the 11B or 90B Vision model has to pin an old engine or move to another model.
llama.cpp added Llama 4 text support in April 2025, and its multimodal documentation now loads Scout with images from a GGUF build (-hf ggml-org/Llama-4-Scout-17B-16E-Instruct-GGUF). It is the route for CPU-heavy servers, Apple silicon and mixed CPU and GPU machines.
Ollama offers both models in its library: Scout at 67 GB in 4-bit, 117 GB in 8-bit and 217 GB at 16-bit, and Maverick at 245 GB, 428 GB and 803 GB. It suits evaluation on a single machine more than a shared production server.
Getting the weights in. On Hugging Face, Meta reviews each access request manually and says approval can take up to a few days. Links from Meta’s own download page expire after 24 hours or five downloads (Meta’s download guide). For an air-gapped site, schedule the request early, record the repository commit and every file’s SHA-256 digest, and move the files in one controlled transfer.
Llama, Qwen, DeepSeek or Mistral: when to pick which
| Family | Licence (verified October 2026) | Sizes worth considering | Pick it when |
|---|---|---|---|
| Llama 4 | Llama 4 Community License; EU exclusion for its multimodal models | Scout 109B and Maverick 400B, both 17B active | You are outside the EU and want a widely supported MoE model with image input and long context |
| Llama 3.3 / 3.1 | Llama 3.x community licences | 8B and 70B dense | You need text only, the broadest tooling, or a Llama model as an EU company |
| Qwen | Apache-2.0 for most checkpoints, including the dense Qwen3.8-27B | Dense and MoE sizes from edge models upward | You want a permissive licence and a size tier for every GPU budget; see our Qwen tier guide |
| DeepSeek V4 | MIT (V4-Flash-0731) | Large MoE checkpoints that need a multi-GPU server | You want frontier-class open weights and can run the DeepSeek provenance review |
| Mistral | Apache-2.0 for Mistral Large 3, Mistral Small 4 and Ministral 3; a modified MIT licence with a large-revenue exception for Mistral Medium 3.5 | Small 4 is 119B with 6.5B active; Large 3 is 675B with 41B active; Ministral 3 comes in 3B, 8B and 14B | You want a European vendor’s models under Apache-2.0, from edge sizes to a single-node flagship |
Mistral’s card recommends FP8 on a single node of B200s or H200s for Large 3, the same hardware class as Maverick. For anything bigger than one node, the Kimi K3 deployment notes show what multi-node serving involves.
How to evaluate a Llama checkpoint before production
- Settle the licence first. Record which entity will run the model, where it is headquartered, and whether any affiliate redistributes products. For an EU company, this step decides between Llama 4 and Llama 3.3.
- Pin and verify the files. Download only from Meta or the
meta-llamaorganisation, record the commit hash, check every safetensors shard’s SHA-256 digest after transfer, and treat community quantizations as separate artifacts that need their own review. - Run your own golden set. Score fifty to a hundred real tasks against the model you would otherwise deploy; launch benchmarks reflect Meta’s settings, not yours.
- Test long context at the length you will use. Plant facts at several depths of a long document and ask for them; Scout’s training context was 256K, well below its advertised limit.
- Check your languages and your dates. Llama 4 lists twelve languages and an August 2024 knowledge cutoff, and Llama 3.3 stops at December 2023, so recent facts must come from retrieval.
- Add guard models. Llama Guard 4 is a 12B multimodal classifier for prompts and responses that runs on a single GPU, and Prompt Guard 2 detects prompt injection and jailbreak attempts in 86M and 22M sizes with a 512-token window. Our comparison of open-weight guard models covers where each check belongs.
- Re-test on every change. A new quantization, engine version or context limit is a new deployment as far as quality is concerned.
How VDF AI fits
VDF AI treats a self-hosted Llama deployment as one more local model. VDF AI Chat lists Llama among the open-weight families it serves inside your perimeter, and VDF AI Agents can use any OpenAI-compatible endpoint, such as a vLLM server running Llama 3.3 70B, as an agent’s model. VDF AI Router applies policy before any learned routing: models can be pinned per workload, organisation-wide allow and deny lists apply to every request, and a domain flagged as regulated only considers approved models. If the licence review rules out Llama 4 for an EU entity, a deny rule keeps it out of every workflow.
Before a Llama checkpoint replaces an incumbent model, the Model Evaluation Suite runs your stored use cases against both inside your deployment, scores each answer against a reference with BLEU, ROUGE-L, METEOR and BERTScore, and keeps every response with timestamps. The choice between Llama, Qwen and Mistral then rests on recorded results from your own tasks.
Sources
- Llama 4 Scout model card and Llama 4 Maverick model card
- Llama 3.3 70B model card, Llama 3.2 Vision model card and Llama 3.1 8B model card
- Meta: the Llama 4 herd
- Llama 4 Community License, Llama 3.3 Community License and Llama 4 acceptable use policy
- Config files (ungated mirrors): Llama 4 Scout, Llama 4 Maverick, Llama 3.3 70B and Llama 3.1 8B
- Meta: getting the models from Meta and from Hugging Face
- Meta: personal superintelligence (July 2025), Muse Spark announcement, Muse Glimmer announcement and Muse Glimmer model card
- Mark Zuckerberg: the future is for everyone (August 2026)
- vLLM: Llama 4 support, supported models and model registry
- llama.cpp: Llama 4 text support and multimodal documentation
- Ollama: llama4 tags
- Llama Guard 4 and Llama Prompt Guard 2
- NVIDIA: LLM inference optimization, H100 and H200
- Qwen3.8-27B, DeepSeek-V4-Flash-0731 and Mistral models overview