Gemma 4 on premise means running Google's Gemma 4 open weights on your own GPUs or laptops. There are five sizes, E2B, E4B, 12B, 26B A4B and 31B, all under Apache 2.0. Local requirements run from about 4 GB for E2B to 25 GB or more for the 31B at 4-bit with a 32K context, and 80 GB-class GPUs for team serving at BF16.
Gemma 4 is the first Gemma generation a legal team can treat like any other Apache 2.0 component. It spans phones and edge boards, a 12B model for 16 GB laptops, and two server-class models that Google says fit one 80 GB H100 at full precision.
This guide covers what to download, what it needs, and how to run it. Sizes, licences and commands were checked against Google’s model cards, the Hugging Face file listings and the engine documentation in October 2026. For the same exercise on other families, see the gpt-oss local setup and the Mistral self-deployment guide.
Gemma 4 quick picks
| If you have | Run | Format | Why |
|---|---|---|---|
| A phone, an edge board or 8 GB of memory | E2B or E4B | Google’s QAT Q4_0 build | 128K context, audio and image input, built for offline devices |
| A 16 GB GPU or laptop | 12B | QAT Q4_0 (about 7 GB) | Google says it nears the 26B MoE on standard benchmarks at under half the memory |
| A 24 GB GPU | 26B A4B | QAT Q4_0 (about 14 GB) | Only 3.8B active parameters per token and a small cache for long sessions |
| A 32 GB GPU | 31B | QAT Q4_0 (about 18 GB) | The strongest Gemma 4 model, with room for a 32K context |
| One 80 GB H100 | 26B A4B at BF16 | Original weights | Ten 32K-token sessions fit with headroom |
| One 141 GB H200 or two H100s | 31B at BF16 | Original weights | Full precision for a team-sized service |
Gemma 4 sizes and variants
Google released Gemma 4 on 2 April 2026 with four sizes and added the 12B model on 3 June 2026. Parameter counts and context windows come from Google’s model card; checkpoint sizes are the safetensors totals on Hugging Face (verified October 2026).
| Size | Architecture | Parameters | Context | Input | BF16 checkpoint |
|---|---|---|---|---|---|
| E2B | Dense with per-layer embeddings | 2.3B effective, 5.1B with embeddings | 128K | Text, image, audio | ≈10.2 GB |
| E4B | Dense with per-layer embeddings | 4.5B effective, 8B with embeddings | 128K | Text, image, audio | ≈16.0 GB |
| 12B | Dense, encoder-free | 11.95B | 256K | Text, image, audio | ≈23.9 GB |
| 26B A4B | Mixture of experts, 128 experts, 8 routed per token | 25.2B total, 3.8B active | 256K | Text, image | ≈51.6 GB |
| 31B | Dense | 30.7B | 256K | Text, image | ≈62.5 GB |
All sizes produce text only, and video arrives as sampled frames. The “E” means effective parameters: the small models carry large per-layer embedding tables that inflate the file without adding compute. The 12B model drops the separate vision and audio encoders and projects raw inputs straight into the language model’s embedding space.
Each size also has a base checkpoint for fine-tuning and a small -it-assistant companion. The Gemma MTP guide describes these as four-layer drafters for speculative decoding: the drafter proposes several tokens and the main model verifies them in one pass, so Google expects similar quality with faster decoding whenever drafts are accepted. The 31B drafter is under 1 GB.
What the Apache 2.0 licence allows
Gemma 4 ships under the plain Apache License 2.0, and none of the Gemma 4 repositories we checked on Hugging Face is gated. Google’s Gemma Terms of Use now state that they cover the models listed in their appendix and send Gemma 4 users to the Apache licence instead. For an enterprise, that changes three things:
- Commercial use needs no request. There is no user threshold, revenue test or acceptance step before download.
- Redistribution is ordinary. Pass on the licence text, mark files you changed, and keep the attribution and NOTICE content. Apache 2.0 also includes a patent licence from contributors.
- No use-restriction flow-down. Section 3.1 of the earlier Gemma terms made you write Google’s use restrictions into any agreement governing a model you distribute. Gemma 4 has no such clause, which makes fine-tuned derivatives easier to share with subsidiaries or customers.
Gemma 1, 2, 3 and 3n, and variants such as ShieldGemma and EmbeddingGemma, stay under the older terms. If your estate mixes generations, track the licence per checkpoint. For EU-headquartered companies, the contrast with Llama 4 is the practical point: Meta withholds its multimodal licence grant from them, as our Llama licence notes explain, while every Gemma 4 size accepts images under Apache 2.0. Our review of clauses that create work covers the remaining checks.
Local requirements: weights, KV cache and headroom
Memory is the sum of weights, KV cache and runtime overhead. Google’s Gemma 4 overview gives the weights side and states that it excludes the context window:
| Size | BF16 | SFP8 | Q4_0 |
|---|---|---|---|
| E2B | 11.4 GB | 5.7 GB | 2.9 GB |
| E4B | 17.9 GB | 8.9 GB | 4.5 GB |
| 12B | 26.7 GB | 13.4 GB | 6.7 GB |
| 26B A4B | 57.7 GB | 28.8 GB | 14.4 GB |
| 31B | 69.9 GB | 34.9 GB | 17.5 GB |
The KV cache depends on the attention layout, read from each config.json. In the 12B, 26B A4B and 31B models, five of every six layers use a 1,024-token sliding window and every sixth layer attends globally with larger heads. An engine that keeps only the window for sliding layers, as vLLM’s hybrid KV cache manager and llama.cpp’s default cache do, pays a fixed amount for those layers and grows only in the global ones. With a 16-bit cache, counting keys and values separately as our VRAM calculator does:
| Size | Layers (sliding / global) | Growth per token | Fixed window cache | One 32K session | One 128K session |
|---|---|---|---|---|---|
| 12B | 40 / 8 | 16 KiB | ≈0.34 GB | ≈0.87 GB | ≈2.5 GB |
| 26B A4B | 25 / 5 | 20 KiB | ≈0.21 GB | ≈0.88 GB | ≈2.9 GB |
| 31B | 50 / 10 | 80 KiB | ≈0.84 GB | ≈3.5 GB | ≈11.6 GB |
The 31B model has ten global layers with four KV heads of dimension 512, against five layers with two such heads in the 26B A4B, so its cache grows four times faster. That gap, more than the weights, separates them for long-context or multi-user serving. The global layers also reuse the key projection as the value (attention_k_eq_v in the config), so an engine that stored it once would need less than shown. E2B and E4B share KV projections across their last 18 to 20 layers, which the standard formula does not model, so size them from Google’s weights table plus a measured margin.
Putting the pieces together, dividing by 0.9 to keep a tenth of memory for activations and the runtime, as in our GPU sizing method:
| Scenario | Weights | KV cache | Needed | Result |
|---|---|---|---|---|
| 12B, Google QAT Q4_0, one 32K session | 7.2 GB | 0.9 GB | ≈8.9 GB | Fits a 12 GB or 16 GB GPU |
| 26B A4B, QAT Q4_0, one 32K session | 15.6 GB | 0.9 GB | ≈18.3 GB | Fits a 24 GB GPU |
| 31B, QAT Q4_0, one 8K session | 18.9 GB | 1.5 GB | ≈22.6 GB | Tight on 24 GB |
| 31B, QAT Q4_0, one 32K session | 18.9 GB | 3.5 GB | ≈24.9 GB | Needs a 32 GB GPU |
| 12B, BF16, ten 32K sessions | 23.9 GB | 8.7 GB | ≈36.3 GB | One 48 GB L40S |
| 26B A4B, BF16, ten 32K sessions | 51.6 GB | 8.8 GB | ≈67.1 GB | One 80 GB H100 |
| 31B, BF16, ten 32K sessions | 62.5 GB | 35.2 GB | ≈108.6 GB | One 141 GB H200 or two H100s |
QAT weights include the image projector file. Ten sessions is what 50 named users give at 20% peak concurrency. The 31B BF16 row explains why the vLLM recipe launches that model across two GPUs.
Google’s quantized builds
Google publishes its own quantization-aware trained (QAT) checkpoints in the google organisation, which settles most provenance questions about 4-bit Gemma (verified October 2026):
- QAT Q4_0 GGUF for all five sizes, for llama.cpp, Ollama and LM Studio. The 31B file is 17.7 GB, the 26B A4B 14.4 GB, the 12B 7.0 GB, E4B 5.2 GB and E2B 3.4 GB, each with a separate projector file for image input. Google says QAT preserves quality similar to BF16 while sharply cutting memory.
- W4A16 compressed-tensors for E2B, E4B, 12B and 31B, made for vLLM. The 31B build is 23.3 GB.
- Unquantized QAT weights (
-qat-q4_0-unquantized) for teams that want to run their own conversion from the QAT-trained starting point. - Mobile QAT builds for E2B and E4B, which the vLLM recipe describes as mixed int2, int4 and int8 compression.
For the 26B A4B on vLLM, where no W4A16 build exists, the vLLM recipe suggests --quantization int8_per_channel_weight_only. Community GGUF and MLX conversions are plentiful, but each is a separate artifact with its own builder; prefer Google’s files unless you have tested the alternative.
Install and run Gemma 4 locally
Pick the engine by where it runs. Our vLLM, Ollama and llama.cpp comparison covers the trade-offs; the commands below are the ones the vendors document.
- Ollama, for a single machine. The library carries every size, with QAT, Q4_K_M, Q8_0, BF16, MLX and NVFP4 tags.
ollama run gemma4:26bpulls the default 26B A4B build;gemma4:31b-it-qat(19 GB) andgemma4:12b-it-qat(7.2 GB) pick Google’s QAT files explicitly. - llama.cpp, for CPU, mixed or Apple hardware. Google’s QAT card gives
llama-server -hf google/gemma-4-31B-it-qat-q4_0-gguf:Q4_0. The server keeps a window-sized cache for sliding layers unless you pass--swa-full. - LM Studio, for desktops. It offers E2B, E4B, 26B A4B and 31B in GGUF and MLX formats, with tool use, vision and reasoning; its catalogue lists about 4 GB of RAM for the smallest size.
- MLX, for Apple silicon. The
mlx-communityconversions usepip install -U mlx-vlm, thenmlx_vlm.generate --model mlx-community/gemma-4-26b-a4b-it-8bit --max-tokens 100. For an OpenAI-compatible endpoint, the card usesmlx_lm.server --model mlx-community/gemma-4-26b-a4b-it-8bit. - vLLM, for shared servers. Install a current build and serve with a bounded context:
vllm serve google/gemma-4-26B-A4B-it --max-model-len 32768 --gpu-memory-utilization 0.90. The recipe runs the 31B with--tensor-parallel-size 2at the same context. - Turn on tools and reasoning in vLLM. Add
--enable-auto-tool-choice --reasoning-parser gemma4 --tool-call-parser gemma4with the recipe’stool_chat_template_gemma4.jinjatemplate, and cap multimodal input with--limit-mm-per-prompt. For speculative decoding, point--speculative-configat the matching-it-assistantdrafter. - Set sampling and thinking. The card recommends temperature 1.0, top_p 0.95 and top_k 64. Thinking is off unless the system prompt begins with the
<|think|>token.
For several GPUs or several models, run vLLM behind a Kubernetes deployment; our Kubernetes hosting notes cover scheduling and GPU allocation.
Before Gemma 4 serves real users
- Pull from the
googleorganisation and pin the commit. Record the repository hash and each file’s SHA-256 digest, and treat any third-party quantization as a separate artifact with its own review. - Block egress on the serving hosts. Set
HF_HUB_OFFLINE=1so the Hugging Face client makes no Hub calls (Hugging Face docs), and opt out of vLLM’s anonymous usage statistics withVLLM_NO_USAGE_STATS=1(vLLM docs). - Decide on audio. E2B, E4B and 12B accept speech directly. Recorded meetings and calls usually contain personal data, so apply your transcript retention rules to the raw audio as well.
- Test the size you will ship. Run your own task set against the QAT build you deploy, not against Google’s BF16 benchmark tables, and repeat it when you change quantization, engine version or context limit.
- Bound the context. Set
--max-model-lento what your workloads use; the 31B model’s cache at full 256K length outweighs its 4-bit weights.
How VDF AI fits
VDF AI treats a Gemma 4 server as another local model. VDF AI Chat and VDF AI Agents accept any OpenAI-compatible endpoint, so a vLLM, Ollama or llama.cpp server running Gemma 4 can back a specific agent while other agents use different models under different policies.
VDF AI Router registers Ollama and custom on-premises deployments next to any cloud models policy allows, probes local runtimes continuously, and has an air-gap mode that restricts routing to local models. That lets a team send short, high-volume requests to the 26B A4B and reserve the 31B for work that needs it. Before a Gemma 4 size replaces an incumbent, the Model Evaluation Suite runs your stored test cases against both inside your deployment and keeps every scored response for the approval record.
Sources
- Google: Gemma 4 launch, 2 April 2026
- Google: Gemma 4 12B launch, 3 June 2026
- Gemma 4 model card
- Gemma 4 overview and memory table
- Gemma 4 licence (Apache 2.0) and Gemma Terms of Use
- Gemma MTP drafters
- Gemma 4 collection on Hugging Face
- gemma-4-31B-it model card and config.json
- Config files: 26B A4B and 12B
- Gemma 4 31B QAT Q4_0 GGUF and W4A16 build
- vLLM recipe: Gemma 4
- vLLM hybrid KV cache manager
- llama.cpp server options
- Ollama: gemma4 tags
- LM Studio: Gemma 4
- MLX community: Gemma 4 26B A4B 8-bit
- Hugging Face Hub environment variables and vLLM usage statistics
- NVIDIA L40S, H100 and H200