gpt-oss local deployment means serving OpenAI's open-weight gpt-oss-20b or gpt-oss-120b on your own hardware. Both are Apache 2.0 mixture-of-experts reasoning models shipped in MXFP4: the 20b needs about 16 GB of memory and the 120b fits one 80 GB GPU. Serve them with vLLM, Ollama or llama.cpp, all of which handle the required harmony format.
gpt-oss is the simplest licence story among the large open models and one of the easiest to size, because OpenAI did the quantization during training. What trips teams up is the response format: the models expect harmony, reason in a separate channel that should not reach end users, and behave differently across the Chat Completions and Responses APIs.
This guide walks through requirements, configuration and setup. Model facts come from OpenAI’s model card, its GitHub repository and the Hugging Face repositories; engine steps come from the vLLM, Ollama and llama.cpp documentation (verified October 2026). Companion guides cover Gemma 4 sizes and Mistral’s open models.
gpt-oss at a glance
| gpt-oss-20b | gpt-oss-120b | |
|---|---|---|
| Released | 5 August 2025 | 5 August 2025 |
| Layers | 24 | 36 |
| Parameters, total / active per token | 20.9B / 3.6B | 116.8B / 5.1B |
| Experts | 32, top 4 per token | 128, top 4 per token |
| Context | 131,072 tokens | 131,072 tokens |
| Weights as published | ≈13.8 GB | ≈65.2 GB |
| OpenAI’s hardware statement | Systems with as little as 16 GB of memory | A single 80 GB GPU |
| Ollama tag | gpt-oss:20b, 14 GB | gpt-oss:120b, 65 GB |
| Sensible starting hardware | A 24 GB GPU, or 16 GB with some experts offloaded to the CPU | One 141 GB H200, or one 80 GB GPU for light use |
Both models are text-only, with a knowledge cutoff of June 2024. They support three reasoning levels, function calling and structured outputs, and were trained to use a browsing tool and a Python tool.
What OpenAI has released, and what it has not
OpenAI published the two models on 5 August 2025 under the Apache 2.0 licence and a gpt-oss usage policy. The policy file in each repository is two sentences long; its operative line asks users to comply with all applicable law. There is no acceptable-use annex, user threshold or naming requirement.
As of October 2026, OpenAI’s Hugging Face organisation lists no newer gpt-oss checkpoint, and the main repositories were last modified in August 2025. The family does have two derivatives: gpt-oss-safeguard-20b and gpt-oss-safeguard-120b, Apache 2.0 models fine-tuned to classify content against a safety policy you write. Our guard model comparison covers where they fit; they are classifiers, not replacements for the base models.
Two architectural details explain the memory profile:
- MXFP4 experts. OpenAI post-trained the models with the mixture-of-experts weights quantized to MXFP4, about 4.25 bits per parameter. Those weights hold over 90% of the parameters, so the checkpoint is roughly a quarter of its BF16 size, and you do not need to choose or validate a quantization yourself.
- Alternating attention. Layers alternate between a 128-token banded window and full attention. Each layer has 64 query heads of dimension 64 and 8 key-value heads, and learned attention sinks let a head attend to nothing.
gpt-oss local requirements by size and context
The KV cache follows from the configs: only the full-attention half of the layers grows with context, and the windowed half holds 128 tokens. At a 16-bit cache that gives 24 KiB per token for the 20b and 36 KiB for the 120b, small next to most dense models. Engines that drop tokens outside the window, as vLLM’s hybrid KV cache manager does for gpt-oss, keep the windowed layers at a few megabytes.
Applying our GPU sizing method, with a tenth of memory held back for activations and the runtime (unified-memory systems keep a fifth for the operating system):
| Scenario | Weights | KV cache | Needed | Result |
|---|---|---|---|---|
| 20b, one 32K session | 13.8 GB | 0.8 GB | ≈16.2 GB | At the limit of a 16 GB GPU |
| 20b, one 128K session | 13.8 GB | 3.2 GB | ≈18.9 GB | Fits a 24 GB GPU |
| 20b, ten 32K sessions | 13.8 GB | 8.1 GB | ≈24.3 GB | Fits a 32 GB GPU |
| 120b, one 128K session | 65.2 GB | 4.8 GB | ≈77.9 GB | Tight on one 80 GB H100 |
| 120b, one 128K session on a 128 GB DGX Spark | 65.2 GB | 4.8 GB | ≈87.6 GB | Fits, with room for a second model |
| 120b, ten 32K sessions | 65.2 GB | 12.1 GB | ≈86.0 GB | Tight on a 94 GB H100 NVL; fits one 141 GB H200 |
| 120b, twenty 32K sessions | 65.2 GB | 24.3 GB | ≈99.4 GB | Fits one H200 at 70% |
Ten sessions is what 50 named users produce at 20% peak concurrency. The llama.cpp maintainers’ gpt-oss guide gives measured totals that agree: 15.5 GB for the 20b and 64.9 GB for the 120b at a 32K context, rising to 17.9 GB and 68.5 GB at 131K. Their figures run lower than ours because they report measured use without the 10% margin we add. Plug other mixes into the VRAM calculator, which has both models as presets.
Harmony, reasoning effort and tool calls
OpenAI states that both models “should only be used with the harmony format.” The harmony guide defines what the format carries:
- Roles in priority order: system, developer, user, assistant, tool. Developer messages hold your instructions and function definitions.
- Channels:
analysisfor chain-of-thought,commentaryfor tool calls, andfinalfor the answer shown to the user. - Reasoning effort: set in the system message as
Reasoning: low,mediumorhigh, with medium as the default. OpenAI’s model card shows accuracy rising with effort, at the cost of longer outputs. - Chain-of-thought across turns: drop earlier reasoning once the model has produced a final answer, but pass it back between tool calls inside the same turn.
The serving engines apply harmony for you, so the practical differences are in the APIs they expose. The vLLM recipe explains that /v1/responses can run built-in tools such as browsing and Python inside the reasoning, while /v1/chat/completions supports function calling only. Ollama exposes Chat Completions at http://localhost:11434/v1 but, per OpenAI’s Ollama guide, does not yet support the Responses API natively. If your agents need built-in tools, plan on vLLM.
Set up gpt-oss with vLLM
vLLM is the route for a shared server with many users.
- Install a current build in a virtual environment:
uv pip install vllm --torch-backend=auto. The vLLM recipe now uses the standard package rather than thevllm==0.10.1+gptosswheel shown on the Hugging Face card. - Download once, serve offline. On a connected staging host, fetch the repository with the Hugging Face CLI (
hf download openai/gpt-oss-120b), record the commit and file digests, and move the files into the serving zone. The repository also holdsoriginal/andmetal/copies for OpenAI’s reference implementations, which a vLLM server does not need. - Serve.
vllm serve openai/gpt-oss-120bstarts the server on one GPU. On Hopper, the recipe adds aGPT-OSS_Hopper.yamlconfig that disables prefix caching and setsmax-num-batched-tokens: 8192; on Blackwell it also sets an FP8 KV cache. - Expect tight memory on one H100. The recipe warns that tensor parallel size 1 on an H100 can run out of memory, and that two-way parallelism there needs
--gpu-memory-utilizationbelow 0.95. - Enable function calling with
--tool-call-parser openai --enable-auto-tool-choice. - Treat built-in tools as integrations. The demo browser tool needs an
EXA_API_KEY, which means calls to an external search API, and OpenAI’s repository calls its reference browser “purely for educational purposes.” The Python tool runs code in a Docker container that OpenAI warns could be problematic under prompt injection. Air-gapped sites should leave both off. - Turn off telemetry: set
HF_HUB_OFFLINE=1(Hugging Face docs) andVLLM_NO_USAGE_STATS=1(vLLM docs).
For a multi-replica deployment, our Kubernetes hosting notes cover GPU scheduling, and the engine comparison explains when vLLM’s batching beats a single-user runtime.
Set up gpt-oss with Ollama or llama.cpp
Ollama is the shortest path on one machine. OpenAI’s guide recommends 16 GB or more of VRAM or unified memory for the 20b and 60 GB or more for the 120b.
ollama pull gpt-oss:20b(orgpt-oss:120b).ollama run gpt-oss:20bfor an interactive session.- Point any OpenAI SDK at
http://localhost:11434/v1. Function calls work through Chat Completions; pass the reasoning back between tool calls as the harmony rules require.
llama.cpp suits mixed CPU and GPU machines and Apple silicon. The maintainers’ guide uses the ggml-org MXFP4 GGUF files, about 12.1 GB and 63.4 GB:
llama-server -hf ggml-org/gpt-oss-20b-GGUF --ctx-size 0 --jinja -ub 2048 -b 2048serves the full context with the built-in chat template.- Set a default reasoning level with
--chat-template-kwargs '{"reasoning_effort": "high"}'. - On a 16 GB NVIDIA card, the guide restricts context to 32K and moves some expert layers to the CPU with
--n-cpu-moe. On 8 GB cards it moves 22 layers for the 20b and 35 for the 120b. - Sample at temperature 1.0 and top_p 1.0, as OpenAI advises, and leave repetition penalties off, as the llama.cpp guide insists.
LM Studio downloads either model with lms get openai/gpt-oss-20b or lms get openai/gpt-oss-120b.
Governance checks from OpenAI’s own model card
The model card is unusually direct about what a deployer must add:
- Keep raw reasoning away from users. OpenAI put no optimisation pressure on the chain-of-thought, so it can contain hallucinated or policy-violating text. The card says developers should not show it to users without filtering, moderation or summarisation.
- Do not rely on the system prompt alone. Both models trail OpenAI o4-mini on OpenAI’s instruction-hierarchy tests, which the card reads as a developer being less able to block jailbreaks through the system message. Put permission checks and approvals outside the model.
- Ground factual answers. Without browsing, OpenAI measured SimpleQA hallucination rates of 78.2% for the 120b and 91.4% for the 20b. Retrieval from your own sources matters more here than with larger hosted models.
- Remember the release cannot be recalled. OpenAI notes that once open weights are out, attackers could fine-tune them to bypass refusals and OpenAI cannot revoke access. Control who can modify and redeploy your copy.
- Record the configuration. Log the checkpoint hash, engine version, reasoning level and enabled tools with each deployment, because each changes behaviour.
How VDF AI fits
VDF AI Chat and VDF AI Agents accept any OpenAI-compatible endpoint, so a vLLM or Ollama server running gpt-oss can serve as an agent’s model while sensitive agents stay on other approved models. VDF AI Router registers those local deployments alongside any cloud models policy allows; organisation-wide allow and deny lists, per-workload pinning and regulated-domain approvals apply before its learned routing, and air-gap mode restricts every request to local models.
When you weigh gpt-oss-20b against gpt-oss-120b, or against a model from another family, the Model Evaluation Suite runs your stored use cases on each inside your deployment, scores the answers against your references and flags regressions between versions. The choice of size then rests on your own results.
Sources
- gpt-oss model card (OpenAI, PDF)
- openai/gpt-oss on GitHub
- gpt-oss-120b on Hugging Face, with config.json and usage policy
- gpt-oss-20b on Hugging Face
- OpenAI on Hugging Face and gpt-oss-safeguard-20b
- OpenAI harmony format guide
- OpenAI: run gpt-oss with Ollama and with vLLM
- vLLM recipe: gpt-oss
- vLLM hybrid KV cache manager and usage statistics
- llama.cpp guide: running gpt-oss and ggml-org GGUF files
- Ollama: gpt-oss tags
- Hugging Face Hub environment variables
- NVIDIA H100, H200 and DGX Spark