AI Infrastructure

Gemma 4 On-Premise: Local Requirements, Sizes and Setup (2026)

Google's Gemma 4 family runs from phone-sized E2B to a 31B dense model, all under Apache 2.0. This guide lists every size and variant, the memory each needs for weights and KV cache, Google's official quantization-aware builds, and the install steps for vLLM, llama.cpp, Ollama, LM Studio and MLX.

Gemma 4 on premise means running Google's Gemma 4 open weights on your own GPUs or laptops. There are five sizes, E2B, E4B, 12B, 26B A4B and 31B, all under Apache 2.0. Local requirements run from about 4 GB for E2B to 25 GB or more for the 31B at 4-bit with a 32K context, and 80 GB-class GPUs for team serving at BF16.

Gemma 4 is the first Gemma generation a legal team can treat like any other Apache 2.0 component. It spans phones and edge boards, a 12B model for 16 GB laptops, and two server-class models that Google says fit one 80 GB H100 at full precision.

This guide covers what to download, what it needs, and how to run it. Sizes, licences and commands were checked against Google’s model cards, the Hugging Face file listings and the engine documentation in October 2026. For the same exercise on other families, see the gpt-oss local setup and the Mistral self-deployment guide.

Gemma 4 quick picks

If you haveRunFormatWhy
A phone, an edge board or 8 GB of memoryE2B or E4BGoogle’s QAT Q4_0 build128K context, audio and image input, built for offline devices
A 16 GB GPU or laptop12BQAT Q4_0 (about 7 GB)Google says it nears the 26B MoE on standard benchmarks at under half the memory
A 24 GB GPU26B A4BQAT Q4_0 (about 14 GB)Only 3.8B active parameters per token and a small cache for long sessions
A 32 GB GPU31BQAT Q4_0 (about 18 GB)The strongest Gemma 4 model, with room for a 32K context
One 80 GB H10026B A4B at BF16Original weightsTen 32K-token sessions fit with headroom
One 141 GB H200 or two H100s31B at BF16Original weightsFull precision for a team-sized service

Gemma 4 sizes and variants

Google released Gemma 4 on 2 April 2026 with four sizes and added the 12B model on 3 June 2026. Parameter counts and context windows come from Google’s model card; checkpoint sizes are the safetensors totals on Hugging Face (verified October 2026).

SizeArchitectureParametersContextInputBF16 checkpoint
E2BDense with per-layer embeddings2.3B effective, 5.1B with embeddings128KText, image, audio≈10.2 GB
E4BDense with per-layer embeddings4.5B effective, 8B with embeddings128KText, image, audio≈16.0 GB
12BDense, encoder-free11.95B256KText, image, audio≈23.9 GB
26B A4BMixture of experts, 128 experts, 8 routed per token25.2B total, 3.8B active256KText, image≈51.6 GB
31BDense30.7B256KText, image≈62.5 GB

All sizes produce text only, and video arrives as sampled frames. The “E” means effective parameters: the small models carry large per-layer embedding tables that inflate the file without adding compute. The 12B model drops the separate vision and audio encoders and projects raw inputs straight into the language model’s embedding space.

Each size also has a base checkpoint for fine-tuning and a small -it-assistant companion. The Gemma MTP guide describes these as four-layer drafters for speculative decoding: the drafter proposes several tokens and the main model verifies them in one pass, so Google expects similar quality with faster decoding whenever drafts are accepted. The 31B drafter is under 1 GB.

What the Apache 2.0 licence allows

Gemma 4 ships under the plain Apache License 2.0, and none of the Gemma 4 repositories we checked on Hugging Face is gated. Google’s Gemma Terms of Use now state that they cover the models listed in their appendix and send Gemma 4 users to the Apache licence instead. For an enterprise, that changes three things:

  • Commercial use needs no request. There is no user threshold, revenue test or acceptance step before download.
  • Redistribution is ordinary. Pass on the licence text, mark files you changed, and keep the attribution and NOTICE content. Apache 2.0 also includes a patent licence from contributors.
  • No use-restriction flow-down. Section 3.1 of the earlier Gemma terms made you write Google’s use restrictions into any agreement governing a model you distribute. Gemma 4 has no such clause, which makes fine-tuned derivatives easier to share with subsidiaries or customers.

Gemma 1, 2, 3 and 3n, and variants such as ShieldGemma and EmbeddingGemma, stay under the older terms. If your estate mixes generations, track the licence per checkpoint. For EU-headquartered companies, the contrast with Llama 4 is the practical point: Meta withholds its multimodal licence grant from them, as our Llama licence notes explain, while every Gemma 4 size accepts images under Apache 2.0. Our review of clauses that create work covers the remaining checks.

Local requirements: weights, KV cache and headroom

Memory is the sum of weights, KV cache and runtime overhead. Google’s Gemma 4 overview gives the weights side and states that it excludes the context window:

SizeBF16SFP8Q4_0
E2B11.4 GB5.7 GB2.9 GB
E4B17.9 GB8.9 GB4.5 GB
12B26.7 GB13.4 GB6.7 GB
26B A4B57.7 GB28.8 GB14.4 GB
31B69.9 GB34.9 GB17.5 GB

The KV cache depends on the attention layout, read from each config.json. In the 12B, 26B A4B and 31B models, five of every six layers use a 1,024-token sliding window and every sixth layer attends globally with larger heads. An engine that keeps only the window for sliding layers, as vLLM’s hybrid KV cache manager and llama.cpp’s default cache do, pays a fixed amount for those layers and grows only in the global ones. With a 16-bit cache, counting keys and values separately as our VRAM calculator does:

SizeLayers (sliding / global)Growth per tokenFixed window cacheOne 32K sessionOne 128K session
12B40 / 816 KiB≈0.34 GB≈0.87 GB≈2.5 GB
26B A4B25 / 520 KiB≈0.21 GB≈0.88 GB≈2.9 GB
31B50 / 1080 KiB≈0.84 GB≈3.5 GB≈11.6 GB

The 31B model has ten global layers with four KV heads of dimension 512, against five layers with two such heads in the 26B A4B, so its cache grows four times faster. That gap, more than the weights, separates them for long-context or multi-user serving. The global layers also reuse the key projection as the value (attention_k_eq_v in the config), so an engine that stored it once would need less than shown. E2B and E4B share KV projections across their last 18 to 20 layers, which the standard formula does not model, so size them from Google’s weights table plus a measured margin.

Putting the pieces together, dividing by 0.9 to keep a tenth of memory for activations and the runtime, as in our GPU sizing method:

ScenarioWeightsKV cacheNeededResult
12B, Google QAT Q4_0, one 32K session7.2 GB0.9 GB≈8.9 GBFits a 12 GB or 16 GB GPU
26B A4B, QAT Q4_0, one 32K session15.6 GB0.9 GB≈18.3 GBFits a 24 GB GPU
31B, QAT Q4_0, one 8K session18.9 GB1.5 GB≈22.6 GBTight on 24 GB
31B, QAT Q4_0, one 32K session18.9 GB3.5 GB≈24.9 GBNeeds a 32 GB GPU
12B, BF16, ten 32K sessions23.9 GB8.7 GB≈36.3 GBOne 48 GB L40S
26B A4B, BF16, ten 32K sessions51.6 GB8.8 GB≈67.1 GBOne 80 GB H100
31B, BF16, ten 32K sessions62.5 GB35.2 GB≈108.6 GBOne 141 GB H200 or two H100s

QAT weights include the image projector file. Ten sessions is what 50 named users give at 20% peak concurrency. The 31B BF16 row explains why the vLLM recipe launches that model across two GPUs.

Google’s quantized builds

Google publishes its own quantization-aware trained (QAT) checkpoints in the google organisation, which settles most provenance questions about 4-bit Gemma (verified October 2026):

  • QAT Q4_0 GGUF for all five sizes, for llama.cpp, Ollama and LM Studio. The 31B file is 17.7 GB, the 26B A4B 14.4 GB, the 12B 7.0 GB, E4B 5.2 GB and E2B 3.4 GB, each with a separate projector file for image input. Google says QAT preserves quality similar to BF16 while sharply cutting memory.
  • W4A16 compressed-tensors for E2B, E4B, 12B and 31B, made for vLLM. The 31B build is 23.3 GB.
  • Unquantized QAT weights (-qat-q4_0-unquantized) for teams that want to run their own conversion from the QAT-trained starting point.
  • Mobile QAT builds for E2B and E4B, which the vLLM recipe describes as mixed int2, int4 and int8 compression.

For the 26B A4B on vLLM, where no W4A16 build exists, the vLLM recipe suggests --quantization int8_per_channel_weight_only. Community GGUF and MLX conversions are plentiful, but each is a separate artifact with its own builder; prefer Google’s files unless you have tested the alternative.

Install and run Gemma 4 locally

Pick the engine by where it runs. Our vLLM, Ollama and llama.cpp comparison covers the trade-offs; the commands below are the ones the vendors document.

  1. Ollama, for a single machine. The library carries every size, with QAT, Q4_K_M, Q8_0, BF16, MLX and NVFP4 tags. ollama run gemma4:26b pulls the default 26B A4B build; gemma4:31b-it-qat (19 GB) and gemma4:12b-it-qat (7.2 GB) pick Google’s QAT files explicitly.
  2. llama.cpp, for CPU, mixed or Apple hardware. Google’s QAT card gives llama-server -hf google/gemma-4-31B-it-qat-q4_0-gguf:Q4_0. The server keeps a window-sized cache for sliding layers unless you pass --swa-full.
  3. LM Studio, for desktops. It offers E2B, E4B, 26B A4B and 31B in GGUF and MLX formats, with tool use, vision and reasoning; its catalogue lists about 4 GB of RAM for the smallest size.
  4. MLX, for Apple silicon. The mlx-community conversions use pip install -U mlx-vlm, then mlx_vlm.generate --model mlx-community/gemma-4-26b-a4b-it-8bit --max-tokens 100. For an OpenAI-compatible endpoint, the card uses mlx_lm.server --model mlx-community/gemma-4-26b-a4b-it-8bit.
  5. vLLM, for shared servers. Install a current build and serve with a bounded context: vllm serve google/gemma-4-26B-A4B-it --max-model-len 32768 --gpu-memory-utilization 0.90. The recipe runs the 31B with --tensor-parallel-size 2 at the same context.
  6. Turn on tools and reasoning in vLLM. Add --enable-auto-tool-choice --reasoning-parser gemma4 --tool-call-parser gemma4 with the recipe’s tool_chat_template_gemma4.jinja template, and cap multimodal input with --limit-mm-per-prompt. For speculative decoding, point --speculative-config at the matching -it-assistant drafter.
  7. Set sampling and thinking. The card recommends temperature 1.0, top_p 0.95 and top_k 64. Thinking is off unless the system prompt begins with the <|think|> token.

For several GPUs or several models, run vLLM behind a Kubernetes deployment; our Kubernetes hosting notes cover scheduling and GPU allocation.

Before Gemma 4 serves real users

  1. Pull from the google organisation and pin the commit. Record the repository hash and each file’s SHA-256 digest, and treat any third-party quantization as a separate artifact with its own review.
  2. Block egress on the serving hosts. Set HF_HUB_OFFLINE=1 so the Hugging Face client makes no Hub calls (Hugging Face docs), and opt out of vLLM’s anonymous usage statistics with VLLM_NO_USAGE_STATS=1 (vLLM docs).
  3. Decide on audio. E2B, E4B and 12B accept speech directly. Recorded meetings and calls usually contain personal data, so apply your transcript retention rules to the raw audio as well.
  4. Test the size you will ship. Run your own task set against the QAT build you deploy, not against Google’s BF16 benchmark tables, and repeat it when you change quantization, engine version or context limit.
  5. Bound the context. Set --max-model-len to what your workloads use; the 31B model’s cache at full 256K length outweighs its 4-bit weights.

How VDF AI fits

VDF AI treats a Gemma 4 server as another local model. VDF AI Chat and VDF AI Agents accept any OpenAI-compatible endpoint, so a vLLM, Ollama or llama.cpp server running Gemma 4 can back a specific agent while other agents use different models under different policies.

VDF AI Router registers Ollama and custom on-premises deployments next to any cloud models policy allows, probes local runtimes continuously, and has an air-gap mode that restricts routing to local models. That lets a team send short, high-volume requests to the 26B A4B and reserve the 31B for work that needs it. Before a Gemma 4 size replaces an incumbent, the Model Evaluation Suite runs your stored test cases against both inside your deployment and keeps every scored response for the approval record.

Sources

Frequently asked questions

What are the Gemma 4 local requirements?

It depends on the size. Google's own table puts the weights alone at about 2.9 GB for E2B, 4.5 GB for E4B, 6.7 GB for 12B, 14.4 GB for 26B A4B and 17.5 GB for 31B in 4-bit Q4_0, roughly four times that in BF16. Add KV cache for your context and a margin for the runtime. In practice a single 32K-token session needs about 9 GB for the 12B, 18 GB for the 26B A4B and 25 GB for the 31B with Google's 4-bit builds.

Can Gemma 4 run on a 16 GB GPU or laptop?

Yes, for the smaller sizes. Google says the 12B model runs locally with 16 GB of VRAM or unified memory, and its official 4-bit build is about 7 GB plus a small projector file. E2B and E4B need even less; LM Studio lists about 4 GB of RAM for the smallest. The 26B A4B and 31B models want 24 to 32 GB once a useful context is added, even with Google's 4-bit builds, so a 16 GB machine is the wrong target for them.

Is Gemma 4 free for commercial use?

Yes. Gemma 4 is released under the standard Apache License 2.0, which allows commercial use, modification and redistribution provided you keep the licence and notices with any copy you pass on. That is a change from Gemma 1 to 3, which remain under Google's custom Gemma Terms of Use and its use restrictions. Store the licence with the weights you approve, and check the terms of any community fine-tune separately, because derivatives can carry their own conditions.

Should I choose Gemma 4 26B A4B or 31B?

Choose the 26B A4B when you serve many users or long contexts. It activates about 3.8B parameters per token, its weights are smaller, and its KV cache grows a quarter as fast as the 31B's. Choose the 31B dense model when quality on hard tasks matters more than throughput and you can give it a 32 GB card or more. Neither accepts audio; for speech input, use the 12B, E4B or E2B models.

How do I turn Gemma 4 thinking mode on or off?

Gemma 4 reasons only when asked. The model card says thinking is enabled by putting the think control token at the start of the system prompt, and disabled by removing it; in Transformers you can pass enable_thinking=False to the chat template. On vLLM, serve with the gemma4 reasoning parser so the reasoning arrives in its own field instead of mixed into the answer. Thinking adds output tokens, so budget latency and context accordingly.

Filed under
Gemmaopen-weight modelslocal LLMquantizationmixture of expertson-premises AI
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading