AI Infrastructure

gpt-oss Local Deployment Guide: Requirements and Setup (2026)

OpenAI's gpt-oss-20b and gpt-oss-120b are Apache 2.0 reasoning models that ship in MXFP4, so they fit one 16 GB or one 80 GB device. This guide covers the memory each size needs at real context lengths, the harmony format and reasoning effort, tool calling, and step-by-step setup with vLLM, Ollama and llama.cpp.

gpt-oss local deployment means serving OpenAI's open-weight gpt-oss-20b or gpt-oss-120b on your own hardware. Both are Apache 2.0 mixture-of-experts reasoning models shipped in MXFP4: the 20b needs about 16 GB of memory and the 120b fits one 80 GB GPU. Serve them with vLLM, Ollama or llama.cpp, all of which handle the required harmony format.

gpt-oss is the simplest licence story among the large open models and one of the easiest to size, because OpenAI did the quantization during training. What trips teams up is the response format: the models expect harmony, reason in a separate channel that should not reach end users, and behave differently across the Chat Completions and Responses APIs.

This guide walks through requirements, configuration and setup. Model facts come from OpenAI’s model card, its GitHub repository and the Hugging Face repositories; engine steps come from the vLLM, Ollama and llama.cpp documentation (verified October 2026). Companion guides cover Gemma 4 sizes and Mistral’s open models.

gpt-oss at a glance

gpt-oss-20bgpt-oss-120b
Released5 August 20255 August 2025
Layers2436
Parameters, total / active per token20.9B / 3.6B116.8B / 5.1B
Experts32, top 4 per token128, top 4 per token
Context131,072 tokens131,072 tokens
Weights as published≈13.8 GB≈65.2 GB
OpenAI’s hardware statementSystems with as little as 16 GB of memoryA single 80 GB GPU
Ollama taggpt-oss:20b, 14 GBgpt-oss:120b, 65 GB
Sensible starting hardwareA 24 GB GPU, or 16 GB with some experts offloaded to the CPUOne 141 GB H200, or one 80 GB GPU for light use

Both models are text-only, with a knowledge cutoff of June 2024. They support three reasoning levels, function calling and structured outputs, and were trained to use a browsing tool and a Python tool.

What OpenAI has released, and what it has not

OpenAI published the two models on 5 August 2025 under the Apache 2.0 licence and a gpt-oss usage policy. The policy file in each repository is two sentences long; its operative line asks users to comply with all applicable law. There is no acceptable-use annex, user threshold or naming requirement.

As of October 2026, OpenAI’s Hugging Face organisation lists no newer gpt-oss checkpoint, and the main repositories were last modified in August 2025. The family does have two derivatives: gpt-oss-safeguard-20b and gpt-oss-safeguard-120b, Apache 2.0 models fine-tuned to classify content against a safety policy you write. Our guard model comparison covers where they fit; they are classifiers, not replacements for the base models.

Two architectural details explain the memory profile:

  • MXFP4 experts. OpenAI post-trained the models with the mixture-of-experts weights quantized to MXFP4, about 4.25 bits per parameter. Those weights hold over 90% of the parameters, so the checkpoint is roughly a quarter of its BF16 size, and you do not need to choose or validate a quantization yourself.
  • Alternating attention. Layers alternate between a 128-token banded window and full attention. Each layer has 64 query heads of dimension 64 and 8 key-value heads, and learned attention sinks let a head attend to nothing.

gpt-oss local requirements by size and context

The KV cache follows from the configs: only the full-attention half of the layers grows with context, and the windowed half holds 128 tokens. At a 16-bit cache that gives 24 KiB per token for the 20b and 36 KiB for the 120b, small next to most dense models. Engines that drop tokens outside the window, as vLLM’s hybrid KV cache manager does for gpt-oss, keep the windowed layers at a few megabytes.

Applying our GPU sizing method, with a tenth of memory held back for activations and the runtime (unified-memory systems keep a fifth for the operating system):

ScenarioWeightsKV cacheNeededResult
20b, one 32K session13.8 GB0.8 GB≈16.2 GBAt the limit of a 16 GB GPU
20b, one 128K session13.8 GB3.2 GB≈18.9 GBFits a 24 GB GPU
20b, ten 32K sessions13.8 GB8.1 GB≈24.3 GBFits a 32 GB GPU
120b, one 128K session65.2 GB4.8 GB≈77.9 GBTight on one 80 GB H100
120b, one 128K session on a 128 GB DGX Spark65.2 GB4.8 GB≈87.6 GBFits, with room for a second model
120b, ten 32K sessions65.2 GB12.1 GB≈86.0 GBTight on a 94 GB H100 NVL; fits one 141 GB H200
120b, twenty 32K sessions65.2 GB24.3 GB≈99.4 GBFits one H200 at 70%

Ten sessions is what 50 named users produce at 20% peak concurrency. The llama.cpp maintainers’ gpt-oss guide gives measured totals that agree: 15.5 GB for the 20b and 64.9 GB for the 120b at a 32K context, rising to 17.9 GB and 68.5 GB at 131K. Their figures run lower than ours because they report measured use without the 10% margin we add. Plug other mixes into the VRAM calculator, which has both models as presets.

Harmony, reasoning effort and tool calls

OpenAI states that both models “should only be used with the harmony format.” The harmony guide defines what the format carries:

  • Roles in priority order: system, developer, user, assistant, tool. Developer messages hold your instructions and function definitions.
  • Channels: analysis for chain-of-thought, commentary for tool calls, and final for the answer shown to the user.
  • Reasoning effort: set in the system message as Reasoning: low, medium or high, with medium as the default. OpenAI’s model card shows accuracy rising with effort, at the cost of longer outputs.
  • Chain-of-thought across turns: drop earlier reasoning once the model has produced a final answer, but pass it back between tool calls inside the same turn.

The serving engines apply harmony for you, so the practical differences are in the APIs they expose. The vLLM recipe explains that /v1/responses can run built-in tools such as browsing and Python inside the reasoning, while /v1/chat/completions supports function calling only. Ollama exposes Chat Completions at http://localhost:11434/v1 but, per OpenAI’s Ollama guide, does not yet support the Responses API natively. If your agents need built-in tools, plan on vLLM.

Set up gpt-oss with vLLM

vLLM is the route for a shared server with many users.

  1. Install a current build in a virtual environment: uv pip install vllm --torch-backend=auto. The vLLM recipe now uses the standard package rather than the vllm==0.10.1+gptoss wheel shown on the Hugging Face card.
  2. Download once, serve offline. On a connected staging host, fetch the repository with the Hugging Face CLI (hf download openai/gpt-oss-120b), record the commit and file digests, and move the files into the serving zone. The repository also holds original/ and metal/ copies for OpenAI’s reference implementations, which a vLLM server does not need.
  3. Serve. vllm serve openai/gpt-oss-120b starts the server on one GPU. On Hopper, the recipe adds a GPT-OSS_Hopper.yaml config that disables prefix caching and sets max-num-batched-tokens: 8192; on Blackwell it also sets an FP8 KV cache.
  4. Expect tight memory on one H100. The recipe warns that tensor parallel size 1 on an H100 can run out of memory, and that two-way parallelism there needs --gpu-memory-utilization below 0.95.
  5. Enable function calling with --tool-call-parser openai --enable-auto-tool-choice.
  6. Treat built-in tools as integrations. The demo browser tool needs an EXA_API_KEY, which means calls to an external search API, and OpenAI’s repository calls its reference browser “purely for educational purposes.” The Python tool runs code in a Docker container that OpenAI warns could be problematic under prompt injection. Air-gapped sites should leave both off.
  7. Turn off telemetry: set HF_HUB_OFFLINE=1 (Hugging Face docs) and VLLM_NO_USAGE_STATS=1 (vLLM docs).

For a multi-replica deployment, our Kubernetes hosting notes cover GPU scheduling, and the engine comparison explains when vLLM’s batching beats a single-user runtime.

Set up gpt-oss with Ollama or llama.cpp

Ollama is the shortest path on one machine. OpenAI’s guide recommends 16 GB or more of VRAM or unified memory for the 20b and 60 GB or more for the 120b.

  1. ollama pull gpt-oss:20b (or gpt-oss:120b).
  2. ollama run gpt-oss:20b for an interactive session.
  3. Point any OpenAI SDK at http://localhost:11434/v1. Function calls work through Chat Completions; pass the reasoning back between tool calls as the harmony rules require.

llama.cpp suits mixed CPU and GPU machines and Apple silicon. The maintainers’ guide uses the ggml-org MXFP4 GGUF files, about 12.1 GB and 63.4 GB:

  1. llama-server -hf ggml-org/gpt-oss-20b-GGUF --ctx-size 0 --jinja -ub 2048 -b 2048 serves the full context with the built-in chat template.
  2. Set a default reasoning level with --chat-template-kwargs '{"reasoning_effort": "high"}'.
  3. On a 16 GB NVIDIA card, the guide restricts context to 32K and moves some expert layers to the CPU with --n-cpu-moe. On 8 GB cards it moves 22 layers for the 20b and 35 for the 120b.
  4. Sample at temperature 1.0 and top_p 1.0, as OpenAI advises, and leave repetition penalties off, as the llama.cpp guide insists.

LM Studio downloads either model with lms get openai/gpt-oss-20b or lms get openai/gpt-oss-120b.

Governance checks from OpenAI’s own model card

The model card is unusually direct about what a deployer must add:

  1. Keep raw reasoning away from users. OpenAI put no optimisation pressure on the chain-of-thought, so it can contain hallucinated or policy-violating text. The card says developers should not show it to users without filtering, moderation or summarisation.
  2. Do not rely on the system prompt alone. Both models trail OpenAI o4-mini on OpenAI’s instruction-hierarchy tests, which the card reads as a developer being less able to block jailbreaks through the system message. Put permission checks and approvals outside the model.
  3. Ground factual answers. Without browsing, OpenAI measured SimpleQA hallucination rates of 78.2% for the 120b and 91.4% for the 20b. Retrieval from your own sources matters more here than with larger hosted models.
  4. Remember the release cannot be recalled. OpenAI notes that once open weights are out, attackers could fine-tune them to bypass refusals and OpenAI cannot revoke access. Control who can modify and redeploy your copy.
  5. Record the configuration. Log the checkpoint hash, engine version, reasoning level and enabled tools with each deployment, because each changes behaviour.

How VDF AI fits

VDF AI Chat and VDF AI Agents accept any OpenAI-compatible endpoint, so a vLLM or Ollama server running gpt-oss can serve as an agent’s model while sensitive agents stay on other approved models. VDF AI Router registers those local deployments alongside any cloud models policy allows; organisation-wide allow and deny lists, per-workload pinning and regulated-domain approvals apply before its learned routing, and air-gap mode restricts every request to local models.

When you weigh gpt-oss-20b against gpt-oss-120b, or against a model from another family, the Model Evaluation Suite runs your stored use cases on each inside your deployment, scores the answers against your references and flags regressions between versions. The choice of size then rests on your own results.

Sources

Frequently asked questions

What are the gpt-oss local requirements?

OpenAI says gpt-oss-20b runs on systems with as little as 16 GB of memory and gpt-oss-120b fits on a single 80 GB GPU, because the expert weights ship in MXFP4. The files are about 13.8 GB and 65.2 GB. Context adds KV cache: roughly 0.8 GB per 32K-token session for the 20b and 1.2 GB for the 120b at 16-bit. A 24 GB card is comfortable for the 20b at full context, and one 141 GB H200 serves the 120b to a team.

Can gpt-oss-120b run on a single GPU?

Yes. OpenAI designed it to fit one 80 GB GPU such as an H100 or MI300X, and the weights take about 65 GB. The margin is small: one 128K-token session brings the total to roughly 78 GB with runtime headroom, and the vLLM recipe warns that a single H100 can run out of memory unless you raise memory utilisation or lower the batched-token setting. For several concurrent users, a 141 GB H200 or two 80 GB GPUs is the safer plan.

Is gpt-oss free for commercial use?

Yes. Both models are released under the Apache 2.0 licence, which allows commercial use, modification and redistribution with the licence and notices kept. Each repository also includes a short gpt-oss usage policy in which OpenAI asks users to comply with all applicable law. There is no user cap, revenue threshold or naming rule. Keep the licence and usage policy files with the weights in your model registry so reviewers can see what was approved.

Why does gpt-oss need the harmony format?

Both models were trained on OpenAI's harmony response format, and OpenAI says they will not work correctly without it. Harmony separates the model's output into channels: analysis for chain-of-thought, commentary for tool calls and final for the answer the user sees. It also carries the reasoning level and the role hierarchy. vLLM, Ollama, llama.cpp and the Transformers chat template apply it for you; problems usually appear when a custom client builds prompts by hand.

Is there a newer gpt-oss model than 20b and 120b?

Not as of October 2026. OpenAI's Hugging Face organisation lists only gpt-oss-20b and gpt-oss-120b, both released in August 2025, plus two gpt-oss-safeguard models that are fine-tuned from them for policy-based safety classification. The main checkpoints have not been updated since late August 2025. Check the organisation page before you pin a version, because a new release would appear there first.

Filed under
gpt-ossopen-weight modelslocal LLMmixture of expertsvLLMOllamaon-premises AI
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading