Sizing tool · presets verified October 2026

LLM VRAM Calculator

GPU memory, KV cache and GPU count for self-hosted LLM inference

An LLM needs GPU memory for two things: its weights, and a KV cache that grows with every token each active conversation holds. Add both, divide by 0.9 for runtime headroom, and you have the memory to buy. At the defaults below, Qwen3-32B at FP8 serving 50 named users with 8,192-token contexts needs about 60.3 GB: one H100 80 GB.

Estimate GPU memory for your model

Pick a preset or type the values from your model’s config.json, then set precision, workload and hardware. The arithmetic runs in your browser and results update as you type.

Every field updates the results. A summary is announced after you pause typing.

Model

Dense, grouped-query attention. 64 layers, 8 KV heads of dimension 128. Context: 32,768 tokens natively, 131,072 with YaRN. Qwen3-32B config.json

Precision

“≈” marks an average that includes each format’s scales.

Workload
Size concurrency from

Teams under 10 people count everyone as active.

Hardware

80 GB, 3.35 TB/s. SXM module on NVLink-connected boards. H100 specs

Advanced settings
Advanced

0.90 keeps a tenth free for activations and the runtime; vLLM’s own default is 0.92.

Embedding, reranker, guard or draft models on the same GPUs.

For runtimes without a sliding-window cache, such as llama.cpp with --swa-full.

Result

60.3 GB of GPU memory needed

Fits Fits on 1 × H100 80 GB, using 75% of its memory.

Model weights
32.8 GB FP8 or INT8 (8 bits per weight)
KV cache, all sequences
21.5 GB
KV cache per token
262,144 bytes
KV cache per sequence
2.1 GB at 8,192 tokens
Concurrent sequences
10 (20% of 50 named users, rounded up)
GPUs or systems
1 × H100 80 GB (auto)
Most sequences at this context
18
Longest context at this concurrency
14,968 tokens
Decode ceiling, one sequence
≈96 tokens/s theoretical upper bound, not a benchmark

Where the memory goes

Memory breakdown for the current settings
ComponentGBShare of GPU memory
Model weights32.8 GB40%
KV cache21.5 GB26%
Other resident models0 GB0%
Headroom held back6.0 GB7%
Unallocated19.7 GB24%
GPU memory in this configuration80 GB100%

How the estimate works

The calculator applies the method from our GPU sizing guide layer by layer, which is what lets it handle models that mix attention types. All sizes are decimal gigabytes (1 GB = 10⁹ bytes), the unit that guide uses, and GPU capacities are compared as the vendors print them.

  1. Weights. Total parameters × bits per weight ÷ 8. A mixture-of-experts model keeps every expert loaded, so the total count applies even though each token touches only the active share.
  2. KV cache per token. In each layer, standard and grouped-query attention store one key and one value per KV head: 2 × KV heads × head dimension values. DeepSeek’s multi-head latent attention stores a single compressed vector instead, kv_lora_rank + qk_rope_head_dim values, which is 576 for DeepSeek-V3.1. Multiply by 2 bytes for a 16-bit cache or 1 byte for FP8, then sum over the layers.
  3. Windowed layers. A sliding-window layer attends only to its most recent tokens and a chunked layer only to its current chunk, so both count min(context, window). vLLM documents that its hybrid KV cache manager reserves just the latest window on sliding-window layers, and it lists Llama 4’s chunked layout among the hybrids it manages; llama.cpp keeps a window-sized cache unless started with --swa-full. If your engine caches these layers at full length, switch on the Advanced toggle.
  4. KV cache in total. KV cache per sequence × concurrent sequences.
  5. Concurrent sequences. Named users × peak-active share, rounded up: 20% by default, 10% for light use. A team of fewer than ten counts everyone. You can also enter sequences directly.
  6. Memory to provision. (Weights + KV cache + other resident models) ÷ usable share. The default share of 0.90 holds a tenth back for activations and the runtime.
  7. GPU count. The smallest of 1, 2, 4 or 8 GPUs whose combined memory covers that figure, then whole 8-GPU nodes. Below 90% of the configuration’s memory the result reads “fits”, from 90% to 100% “tight”, and above that “does not fit”.
  8. Decode ceiling. Memory bandwidth ÷ (bytes of active weights + one sequence’s KV cache). Each generated token reads both once, so this caps single-stream speed from above; real throughput lands lower.

To check the arithmetic by hand, take Qwen3-8B’s config.json: 2 × 36 layers × 8 KV heads × 128 dimensions × 2 bytes comes to 147,456 bytes per token, so one 8,192-token conversation holds 1.2 GB of cache.

Bits per weight by format

Quantized formats store a scale next to each block of weights, so their real cost sits above the nominal bit width. The GGUF figures are the averages llama.cpp measured on Llama 3.1 8B. For what each format does to accuracy, see how quantization formats differ.

FormatBits per weightWhere the figure comes from
BF16 / FP1616Two bytes per weight, as most checkpoints are published.
FP8 or INT88One byte per weight; per-channel or per-block scales add a negligible amount.
AWQ / GPTQ 4-bit≈4.254-bit integers plus a 16-bit scale for each group of weights.
MXFP4≈4.25FP4 values in blocks of 32 that share one 8-bit scale.
NVFP4≈4.5FP4 values in blocks of 16 that share one FP8 scale.
GGUF Q8_0≈8.5llama.cpp measures 8.50 bits per weight.
GGUF Q6_K≈6.6llama.cpp measures 6.56 bits per weight.
GGUF Q5_K_M≈5.7llama.cpp measures 5.70 bits per weight.
GGUF Q4_K_M≈4.9llama.cpp measures 4.89 bits per weight.

VRAM by model and context

The table shows the GPU memory one conversation needs, with a 16-bit KV cache and the 10% reserve included. It was computed when this page was built, by the same module the calculator runs. For several concurrent users, add one more conversation’s KV cache per extra sequence, or let the calculator do it.

GB needed for one sequence, 16-bit KV cache, usable share 0.90
Model Parameters FP8, 8K FP8, 32K 4-bit, 8K 4-bit, 32K
Qwen3-8B 8.2B 10.4 GB14.5 GB6.2 GB10.2 GB
Qwen3-32B 32.8B 38.8 GB45.9 GB21.7 GB28.9 GB
Devstral Small 2 (24B) 24B 28.2 GB32.6 GB15.7 GB20.1 GB
Gemma 4 31B 31.3B 36.4 GB38.7 GB20.1 GB22.4 GB
Llama 3.3 70B 70.6B 81.4 GB90.3 GB44.6 GB53.6 GB
gpt-oss-20b 20.9B total, 3.6B active 23.5 GB24.1 GB12.6 GB13.2 GB
gpt-oss-120b 117B total, 5.1B active 130 GB131 GB69.3 GB70.3 GB
Llama 4 Scout (17B × 16E) 109B total, 17B active 123 GB124 GB65.9 GB67.3 GB
Llama 4 Maverick (17B × 128E) 402B total, 17B active 448 GB449 GB239 GB240 GB
DeepSeek-V3.1 (671B) 671B total, 37B active 746 GB748 GB397 GB399 GB
  • 8K is 8,192 tokens and 32K is 32,768. 4-bit means AWQ or GPTQ at ≈4.25 bits per weight; a GGUF Q4_K_M file runs about 15% larger.
  • gpt-oss ships in MXFP4, while Devstral Small 2 and DeepSeek-V3.1 ship in FP8. Pick “As published” in the calculator to use those exact file sizes.
  • From 8K to 32K, the cache per conversation grows by 8.1 GB on Llama 3.3 70B, where every layer attends to the whole context, but by only 2.0 GB on Gemma 4 31B and 0.9 GB on gpt-oss-120b, whose sliding-window layers stop growing.

Choosing which open-weight model to run comes first; our local LLM overview covers the 2026 options and what each one is good at.

From headcount to GPUs

Capacity requests tend to arrive as a number of employees rather than a model name. Only a slice of them has a request in flight at the busiest moment, so the calculator converts named users into concurrent sequences first: 20% unless you choose 10%, and everyone in a team smaller than ten. Our sizing walkthrough for 5, 50 and 200 users explains those assumptions; the table recomputes its scenarios with this page’s presets.

Worked scenarios at a 20% peak share and usable share 0.90
Named users Sequences Model, weights / KV cache Context Needed Hardware (auto) Status
5 5 Qwen3-32B, FP8 / 16-bit KV 16,384 60.3 GB 1 × H100 80 GB Fits (75%)
5 5 Qwen3-32B, 4-bit AWQ / FP8 KV 16,384 31.3 GB 1 × L40S 48 GB Fits (65%)
50 10 Llama 3.3 70B, 4-bit AWQ / 16-bit KV 8,192 71.5 GB 1 × H100 80 GB Fits (89%)
50 10 Llama 3.3 70B, FP8 / 16-bit KV 8,192 108 GB 1 × H200 141 GB Fits (76%)
200 40 Qwen3-8B, BF16 / 16-bit KV 8,192 71.9 GB 1 × H100 80 GB Fits (89%)
200 40 Qwen3-8B, BF16 / FP8 KV 8,192 45.0 GB 1 × L40S 48 GB Tight (93%)
200 40 Llama 3.3 70B, FP8 / 16-bit KV 8,192 198 GB 2 × H200 141 GB Fits (70%)
200 40 Llama 3.3 70B, FP8 / FP8 KV 8,192 138 GB 1 × H200 141 GB Tight (97%)

The guide rounds its 70B model to 70 billion parameters, while the Llama 3.3 preset carries the published 70.55 billion. That is why the two 200-user FP8 rows read 198 GB and 138 GB here, about 1 GB above the guide’s 197 and 137 GB.

A count that only just fits leaves no room for a failed GPU or a node taken down for patching, so production plans usually add one more server than the arithmetic asks for. The reference architecture lists GPU server classes by model size, and the server buying guide covers power, cooling and the software stack around these GPUs.

GPU memory and bandwidth

Memory capacity decides what fits; memory bandwidth decides how quickly one conversation can generate. The presets use the figures on each vendor’s product page (verified October 2026).

Accelerator presets and their sources
Accelerator Memory Bandwidth Notes Source
NVIDIA L40S 48 GB GDDR6 with ECC 864 GB/s PCIe card without NVLink, so multi-GPU traffic crosses the PCIe bus. L40S specs
NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB GDDR7 1,597 GB/s PCIe card for standard rack servers. RTX PRO 6000 specs
NVIDIA H100 SXM 80 GB 3.35 TB/s SXM module on NVLink-connected boards. H100 specs
NVIDIA H100 NVL 94 GB 3.9 TB/s PCIe version of the H100. H100 NVL specs
NVIDIA H200 (SXM or NVL) 141 GB 4.8 TB/s SXM and NVL versions list the same memory and bandwidth. H200 specs
NVIDIA B200 (one GPU of a DGX B200) 180 GB HBM3e 8 TB/s NVIDIA lists DGX B200 totals (1,440 GB and 64 TB/s across eight GPUs); the per-GPU figures divide them by eight. DGX B200 specs
AMD Instinct MI300X 192 GB HBM3 5.3 TB/s Runs vLLM through ROCm. MI300X specs
AMD Instinct MI325X 256 GB HBM3E 6 TB/s Same platform family as the MI300X, with more and faster memory. MI325X specs
AMD Instinct MI355X 288 GB HBM3E 8 TB/s Direct liquid-cooled. MI355X specs
NVIDIA DGX Spark (128 GB unified) 128 GB LPDDR5x, unified 273 GB/s Unified memory shared with the operating system. Up to three units cable directly and four through a switch. DGX Spark specs
Apple Mac Studio, M5 Max with 128 GB 128 GB unified 614 GB/s The 128 GB option needs the 40-core GPU, which brings 614 GB/s. Memory is shared with macOS. Mac Studio specs
Apple Mac Studio, M5 Ultra with 256 GB 256 GB unified 1.2 TB/s Requires the 36-core CPU and 80-core GPU configuration. Memory is shared with macOS. Mac Studio specs
Apple Mac Studio, M5 Ultra with 512 GB 512 GB unified 1.2 TB/s Listed as a configurable option on the spec page; the bandwidth matches the 256 GB model. Mac Studio specs
AMD Ryzen AI Max+ 395 system with 128 GB 128 GB LPDDR5x-8000, unified 256 GB/s Figures from AMD’s Ryzen AI Halo system. How much of the pool the GPU may address depends on the system and OS. Ryzen AI Halo specs

Desk-side systems give the operating system, the runtime and the model one shared pool, so selecting one sets the usable share to 0.80. The calculator links at most four of them, the most NVIDIA documents for DGX Spark through a switch. Linked systems split a model by layers, which adds memory but no single-stream bandwidth; the desk-side hardware comparison covers the trade-offs between them.

What the calculator leaves out

A memory estimate tells you whether a plan is plausible before you buy, not that a model will load on the first try. These effects sit inside the reserve or on top of it:

  • Activations and prefill buffers. Processing long prompts in large batches needs working memory that grows with the batch’s token count; the reserve covers ordinary batches, not every prefill setting.
  • Fragmentation. Paged KV caches hand out memory in fixed-size blocks, so a little is stranded at the end of each sequence.
  • CUDA graphs and runtime state. The CUDA or ROCm context, communication buffers and captured CUDA graphs claim memory on every GPU before the first request arrives.
  • Multi-GPU and multi-node overhead. Tensor parallelism shards the weights but adds communication buffers, and pipeline stages spread across nodes bring their own buffers and slower links.
  • Speculative decoding. A draft model or multi-token-prediction head adds weights and a KV cache of its own; enter it under other resident models.
  • Shared prefixes. Prefix caching can store a long system prompt once for many conversations; the calculator assumes nothing is shared.
  • Hybrid linear-attention models. Layers built on linear attention or state-space blocks keep a fixed-size state per sequence rather than a growing cache. Model only the full-attention layers as a custom entry and add the state as another resident model.
  • Speed. The decode ceiling ignores compute, batching and interconnects, and prefill is limited by compute rather than bandwidth. For mixed-precision checkpoints it also spreads the file’s average bits over the active parameters, so it can run high. Measure throughput with your own prompts.

Run it on your own servers

Once the GPUs are sized, VDF AI runs on them as the platform layer, whether on-premises, in a private cloud or air-gapped. Its model router can send each request to a local Ollama or other on-prem deployment, and an air-gap mode allows local models only.

LLM VRAM questions

How much VRAM do I need to run a 70B model?

Llama 3.3 70B has 70.55 billion parameters, so its weights take about 141 GB at 16-bit, 70.6 GB at FP8, 37.5 GB at 4-bit AWQ or GPTQ and 43.1 GB as a GGUF Q4_K_M file. Each 8,192-token conversation adds 2.7 GB of 16-bit KV cache. With one conversation and a 10% reserve, plan on 44.6 GB at 4-bit, which is tight on a 48 GB L40S, or 81.4 GB at FP8, which needs a 94–96 GB card such as an H100 NVL or RTX PRO 6000, or a 141 GB H200. Every extra concurrent conversation adds another 2.7 GB.

How do I calculate KV cache size?

Multiply 2 (one key and one value) by the number of layers, the KV heads, the head dimension and the bytes per value: 2 for a 16-bit cache, 1 for FP8. The three shape numbers are num_hidden_layers, num_key_value_heads and head_dim in the model’s config.json. Qwen3-8B gives 2 × 36 × 8 × 128 × 2 = 147,456 bytes per token, about 1.2 GB for an 8,192-token conversation. Multiply by context length and by concurrent sequences. DeepSeek-style MLA layers store kv_lora_rank + qk_rope_head_dim values instead, and sliding-window layers stop growing at their window.

How many GPUs do I need for LLM inference?

Add the weights, the KV cache for your busiest moment and any other models sharing the cards, divide by 0.9 so the runtime keeps a tenth of memory free, then step up through 1, 2, 4 and 8 GPUs until the memory covers it. A 50-person team with a fifth of its members busy at once, running Llama 3.3 70B at FP8 with 8,192-token contexts, comes to about 108 GB. A single 141 GB H200 holds that, as does a pair of 80–96 GB cards splitting the model between them. Memory sets the count; a load test with your own prompts confirms the speed.

Does quantization reduce VRAM requirements?

It shrinks the weights, not the KV cache. FP8 halves 16-bit weights, and 4-bit formats cut them to roughly a quarter once their scales are counted: about 4.25 bits per weight for AWQ, GPTQ and MXFP4, 4.5 for NVFP4 and 4.89 for GGUF Q4_K_M. The KV cache stays at 16-bit unless you quantize it separately, for example with vLLM’s --kv-cache-dtype fp8, which halves it. Check answer quality on your own evaluation set after either change, because the effect differs by model and task.

Can I run a local LLM on a Mac Studio or DGX Spark?

Yes, within the limits of memory and bandwidth. DGX Spark and Ryzen AI Max+ 395 systems have 128 GB of unified memory at 273 and 256 GB/s, and a Mac Studio with M5 Ultra offers up to 512 GB at 1.2 TB/s. The operating system shares that pool, so the calculator counts 80% of it as usable. gpt-oss-120b as published needs about 83.1 GB with one 32,768-token conversation, which fits any of them. Bandwidth then caps generation speed, and every extra user’s KV cache comes out of the same pool.

Why does context length change VRAM so much?

The KV cache grows with every token each active conversation keeps, while the weights stay the same size. Forty concurrent 32,768-token conversations on Llama 3.3 70B need about 429 GB of 16-bit KV cache, roughly six times the 70.6 GB of FP8 weights. Models with windowed attention grow more slowly on engines that cache only the window: gpt-oss and Gemma 4 limit their sliding-window layers to 128 and 1,024 tokens, and Llama 4 limits three of every four layers to 8,192-token chunks. Size the context to what the workload uses rather than the model’s maximum.