AI Infrastructure

Llama On-Premise: Licence, GPU Sizing and Serving for Llama 4 and Llama 3.3

Meta has not released a new Llama model since April 2025, yet Llama 4 Scout, Llama 4 Maverick and Llama 3.3 70B remain common choices for self-hosting. This guide covers the current lineup, the licence terms that bind an enterprise (including the EU carve-out for multimodal models), GPU sizing from the published configs, serving with vLLM, llama.cpp or Ollama, and when another open family is the better fit.

Llama on premise means running Meta's open-weight Llama models on hardware you control. As of October 2026 the current releases are Llama 4 Scout and Maverick, both from April 2025, and Llama 3.3 70B. Meta's community licences allow commercial self-hosting but add a 700-million-user threshold, naming and attribution rules, an acceptable use policy and an EU exclusion for multimodal models.

Llama is often the first open-weight family an enterprise tries, because every serving engine and most tooling support it. Three things still need a decision: licence clauses that other families lack, a Llama 4 design that trades memory for speed, and a line-up that has not changed in eighteen months. Facts below were checked on Hugging Face, Meta’s developer site and the vLLM documentation on 2 October 2026.

The Llama family as of October 2026

ModelReleasedParametersContextInput
Llama 4 Scout5 Apr 2025109B total, 17B active, 16 experts10M tokensText and images
Llama 4 Maverick5 Apr 2025400B total, 17B active, 128 experts1M tokensText and images
Llama 3.3 70B6 Dec 202470B dense128K tokensText
Llama 3.225 Sep 20241B and 3B text; 11B and 90B Vision128K tokensText; text and images for Vision
Llama 3.123 Jul 20248B, 70B and 405B dense128K tokensText

Both Llama 4 models have a knowledge cutoff of August 2024 and list twelve supported languages, among them English, French, German, Italian, Portuguese and Spanish. Meta’s launch post says Scout was pre-trained and post-trained with a 256K context length, so the 10M-token figure relies on length generalisation. Test the context you actually need before you plan around it.

Llama 4 Behemoth was described as still training at launch, and no Behemoth checkpoint has appeared in Meta’s Hugging Face organisation (verified October 2026). In July 2025 Mark Zuckerberg wrote that Meta would need to be careful about what it chooses to open source (Meta). Meta’s newer models carry the Muse name: Muse Spark launched in April 2026 without open weights, offered to developers through Meta’s API, and Muse Glimmer, released on 10 August 2026, is a 30B open-weight model under the Apache 2.0 licence, aimed at local agents on a single consumer GPU. The same day, Zuckerberg wrote that Meta “will resume releasing some open source models soon” (essay).

For planning, treat Llama 4 and Llama 3.3 as the Llama-branded options, and Muse Glimmer as a separate family with different licence terms.

What the Llama licence asks of an enterprise

Each Llama version has its own community licence and acceptable use policy. We read the Llama 4 licence, the Llama 3.3 licence and the Llama 4 acceptable use policy on 2 October 2026:

ClauseWhat it saysWhat it means on premises
Acceptable useThe acceptable use policy is incorporated by referenceFold its prohibited uses into your own AI acceptable-use policy
700 million usersIf your products had more than 700 million monthly active users in the month before the version’s release date, you must request a licence from Meta, which it may grant at its sole discretionIrrelevant for almost every enterprise, but it is counted across affiliates
RedistributionShip a copy of the agreement, display “Built with Llama” prominently, start the names of derived models with “Llama” and include Meta’s attribution noticeApplies when you distribute Llama or a product built on it; sharing a fine-tuned model with another legal entity is likely to count
EU and multimodal modelsFor the multimodal models in Llama 4, the Section 1(a) rights are not granted to individuals domiciled in the EU or companies with their principal place of business thereExcludes EU-headquartered companies from self-hosting Scout and Maverick under the standard grant

The EU clause needs the most attention. Both Llama 4 models accept images, so both are multimodal, and the exemption covers end users of a product that incorporates them, not a company running the weights itself. Llama 3.2’s policy has the same clause for its 11B and 90B Vision models, while Llama 3.3 70B and the text-only 3.1 and 3.2 models are outside it. An EU-headquartered organisation that wants Llama on its own GPUs is therefore choosing between Llama 3.3 70B and older text models; counsel should confirm how the clause applies to your group structure.

How Llama 4 uses memory

Two design choices separate Llama 4 from Llama 3.3. The figures come from the config.json files, read from ungated mirrors because Meta’s own repositories require approved access.

Mixture of experts. Scout has 16 experts in every one of its 48 layers; Maverick has 128 experts in alternating layers, 24 of the 48. Each token is routed to one expert, so both models compute with about 17B active parameters while keeping all 109B or 400B in GPU memory. A mixture-of-experts model therefore generates text at roughly the cost of a much smaller dense model, while its weights cost the full amount.

Mixed attention. Both models use 40 query heads and 8 key-value heads of dimension 128. In three of every four layers, attention is chunked: a token attends only within its 8,192-token chunk, using rotary position embeddings. Every fourth layer, 12 of the 48, attends across the whole context without positional encoding. Because the chunked layers never look past the current chunk, an engine only has to keep that chunk’s keys and values for them, and the config’s cache_implementation: hybrid setting reflects this.

Llama 3.3 70B is a conventional dense model: 80 layers, 64 query heads and 8 key-value heads of dimension 128, with full attention in every layer, so its KV cache grows by the same amount for every token at any context length.

Sizing GPUs for Llama

Two numbers decide the fit. Weights take roughly parameters × bytes per parameter. The KV cache takes 2 × layers × KV heads × head dimension × bytes per value for every token of every concurrent sequence, as in NVIDIA’s derivation and our GPU sizing method. The table applies that to the published configs at a 16-bit KV cache:

ModelWeights: BF16 / FP8 / 4-bitKV cache per tokenPer 8,192-token sequencePer 131,072-token sequence
Llama 3.1 8B≈16 / ≈8 / ≈5 GB128 KiB≈1.1 GB≈17 GB
Llama 3.3 70B≈141 / ≈71 / ≈35–40 GB320 KiB≈2.7 GB≈43 GB
Llama 4 Scout≈217 / ≈109 / ≈55–67 GB192 KiB; 48 KiB per token past the first 8,192 with a hybrid cache≈1.6 GB≈7.7 GB hybrid, ≈26 GB if every layer is cached in full
Llama 4 Maverick≈803 / ≈417 / ≈200–245 GBSame as Scout≈1.6 GBSame as Scout

The BF16 sizes for Llama 4 and the FP8 size for Maverick are the totals in the safetensors index files; Meta publishes no FP8 build of Scout, so its FP8 figure is arithmetic. The 4-bit ranges run from half a byte per parameter up to the size of Ollama’s 4-bit builds, 67 GB for Scout and 245 GB for Maverick. GPU capacities are NVIDIA’s: 80 GB for an H100 SXM and 141 GB for an H200.

Worked example. Suppose 100 to 200 named users, which at 10–20% concurrency gives about 20 sequences at once, each with a 16,384-token context. Divide the total by 0.9 to leave a tenth of GPU memory for activations and the runtime.

  • Llama 3.3 70B at FP8: 71 GB of weights plus 20 × 5.4 GB of KV cache is about 178 GB, or 198 GB with headroom. That is two H200s or four H100s. An FP8 KV cache brings it down to about 138 GB: one H200 at its limit, or two H100s.
  • Llama 4 Scout at BF16, hybrid cache: each sequence holds about 0.8 GB for the global layers and 1.2 GB for one chunk, so 20 sequences need about 40 GB. With 217 GB of weights the total is about 286 GB with headroom, which means four H100s.
  • Llama 4 Scout at 4-bit: around 58 GB of weights plus the same 40 GB of KV cache comes to about 109 GB with headroom: one H200, or two H100s.

Scout needs more memory to load but much less per long conversation, and it computes with 17B parameters per token against 70B. Llama 3.3 is cheaper to hold and simpler to serve. For the server around the GPUs, see the on-premise AI server guide.

Serving Llama with vLLM, llama.cpp and Ollama

vLLM lists Llama4ForConditionalGeneration among its supported models, with text and image input (supported models). Its Llama 4 launch notes required vLLM 0.8.3 or later and reported these limits:

  • On eight H100s, Scout with a 1M-token context and Maverick with about 430K.
  • On eight H200s, Scout up to 3.6M tokens and Maverick up to 1M.
  • The commands used --tensor-parallel-size 8 and --max-model-len, served Maverick from its FP8 checkpoint, and set attn_temperature_tuning through --override-generation-config for Scout’s long-context run.
  • --kv-cache-dtype fp8 can roughly double the usable context, at a quality cost you should measure.

One retirement matters for older deployments: vLLM’s model registry records that the Llama 3.2 Vision architecture “was supported in vLLM until v0.10.2”. A team still running the 11B or 90B Vision model has to pin an old engine or move to another model.

llama.cpp added Llama 4 text support in April 2025, and its multimodal documentation now loads Scout with images from a GGUF build (-hf ggml-org/Llama-4-Scout-17B-16E-Instruct-GGUF). It is the route for CPU-heavy servers, Apple silicon and mixed CPU and GPU machines.

Ollama offers both models in its library: Scout at 67 GB in 4-bit, 117 GB in 8-bit and 217 GB at 16-bit, and Maverick at 245 GB, 428 GB and 803 GB. It suits evaluation on a single machine more than a shared production server.

Getting the weights in. On Hugging Face, Meta reviews each access request manually and says approval can take up to a few days. Links from Meta’s own download page expire after 24 hours or five downloads (Meta’s download guide). For an air-gapped site, schedule the request early, record the repository commit and every file’s SHA-256 digest, and move the files in one controlled transfer.

Llama, Qwen, DeepSeek or Mistral: when to pick which

FamilyLicence (verified October 2026)Sizes worth consideringPick it when
Llama 4Llama 4 Community License; EU exclusion for its multimodal modelsScout 109B and Maverick 400B, both 17B activeYou are outside the EU and want a widely supported MoE model with image input and long context
Llama 3.3 / 3.1Llama 3.x community licences8B and 70B denseYou need text only, the broadest tooling, or a Llama model as an EU company
QwenApache-2.0 for most checkpoints, including the dense Qwen3.8-27BDense and MoE sizes from edge models upwardYou want a permissive licence and a size tier for every GPU budget; see our Qwen tier guide
DeepSeek V4MIT (V4-Flash-0731)Large MoE checkpoints that need a multi-GPU serverYou want frontier-class open weights and can run the DeepSeek provenance review
MistralApache-2.0 for Mistral Large 3, Mistral Small 4 and Ministral 3; a modified MIT licence with a large-revenue exception for Mistral Medium 3.5Small 4 is 119B with 6.5B active; Large 3 is 675B with 41B active; Ministral 3 comes in 3B, 8B and 14BYou want a European vendor’s models under Apache-2.0, from edge sizes to a single-node flagship

Mistral’s card recommends FP8 on a single node of B200s or H200s for Large 3, the same hardware class as Maverick. For anything bigger than one node, the Kimi K3 deployment notes show what multi-node serving involves.

How to evaluate a Llama checkpoint before production

  1. Settle the licence first. Record which entity will run the model, where it is headquartered, and whether any affiliate redistributes products. For an EU company, this step decides between Llama 4 and Llama 3.3.
  2. Pin and verify the files. Download only from Meta or the meta-llama organisation, record the commit hash, check every safetensors shard’s SHA-256 digest after transfer, and treat community quantizations as separate artifacts that need their own review.
  3. Run your own golden set. Score fifty to a hundred real tasks against the model you would otherwise deploy; launch benchmarks reflect Meta’s settings, not yours.
  4. Test long context at the length you will use. Plant facts at several depths of a long document and ask for them; Scout’s training context was 256K, well below its advertised limit.
  5. Check your languages and your dates. Llama 4 lists twelve languages and an August 2024 knowledge cutoff, and Llama 3.3 stops at December 2023, so recent facts must come from retrieval.
  6. Add guard models. Llama Guard 4 is a 12B multimodal classifier for prompts and responses that runs on a single GPU, and Prompt Guard 2 detects prompt injection and jailbreak attempts in 86M and 22M sizes with a 512-token window. Our comparison of open-weight guard models covers where each check belongs.
  7. Re-test on every change. A new quantization, engine version or context limit is a new deployment as far as quality is concerned.

How VDF AI fits

VDF AI treats a self-hosted Llama deployment as one more local model. VDF AI Chat lists Llama among the open-weight families it serves inside your perimeter, and VDF AI Agents can use any OpenAI-compatible endpoint, such as a vLLM server running Llama 3.3 70B, as an agent’s model. VDF AI Router applies policy before any learned routing: models can be pinned per workload, organisation-wide allow and deny lists apply to every request, and a domain flagged as regulated only considers approved models. If the licence review rules out Llama 4 for an EU entity, a deny rule keeps it out of every workflow.

Before a Llama checkpoint replaces an incumbent model, the Model Evaluation Suite runs your stored use cases against both inside your deployment, scores each answer against a reference with BLEU, ROUGE-L, METEOR and BERTScore, and keeps every response with timestamps. The choice between Llama, Qwen and Mistral then rests on recorded results from your own tasks.

Sources

Frequently asked questions

Can you run Llama on premise?

Yes. Meta publishes the Llama weights for download from Hugging Face, after a manual access review, or from its own download page, and the community licence allows commercial self-hosting. Llama 3.3 70B runs on one or two data-centre GPUs depending on precision. Llama 4 Scout fits one 80 GB H100 with 4-bit quantization, according to Meta, while Maverick needs an eight-GPU server. vLLM, llama.cpp and Ollama all serve the Llama 4 models.

Is Llama free for commercial use?

Mostly, yes. There is no licence fee, but the community licence attaches conditions. Organisations whose products had more than 700 million monthly active users in the month before a version's release date must request a separate licence from Meta. Every user must follow Meta's acceptable use policy. Anyone who redistributes Llama or a product built on it must include the licence, display Built with Llama and start the names of derived models with Llama. Llama 4 adds an EU exclusion for its multimodal models.

Can companies in the EU use Llama 4?

Not under the standard licence grant, read literally. Llama 4's acceptable use policy says the rights in Section 1(a) of the licence are not granted for its multimodal models to individuals domiciled in the EU or companies with their principal place of business there, and Scout and Maverick both accept images. End users of a product that incorporates the models are exempt. Llama 3.3 70B is text-only and outside that clause. Have counsel review your group structure before deciding.

How much GPU memory does Llama 4 Scout need?

The BF16 checkpoint is about 217 GB, so at full precision it needs four 80 GB H100s or two 141 GB H200s before any KV cache. Meta says Scout fits on a single H100 with on-the-fly int4 quantization, and Ollama's 4-bit build is 67 GB. Its KV cache is small for its size: 192 KiB per token at 16-bit, falling to 48 KiB per token beyond the first 8,192 tokens when the engine keeps only the current chunk for its local-attention layers.

Should we choose Llama 3.3 70B or Llama 4 Scout?

Llama 3.3 70B has smaller weights, about 141 GB at BF16 against 217 GB for Scout, and is text-only, which keeps it outside the EU multimodal clause. Scout activates only 17B parameters per token, accepts images and keeps a much smaller KV cache for long conversations, so it can serve more long sessions per gigabyte once its weights are loaded. Test both on your own tasks; the published benchmarks will not match your workload.

Is there a newer Llama model than Llama 4?

No new Llama model has appeared since Llama 4 Scout and Maverick in April 2025, and Llama 4 Behemoth, announced as still training, has not been released (verified October 2026). Meta's newer models carry the Muse name. Muse Spark is available only through Meta's API. Muse Glimmer, released in August 2026, is a 30B open-weight model under Apache 2.0. In August 2026 Mark Zuckerberg wrote that Meta will resume releasing some open source models soon.

Filed under
Llamaopen-weight modelslocal LLMmixture of expertson-premises AIAI infrastructure
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading