Speed estimate · GPU specs verified October 2026

LLM Tokens per Second Calculator

Decode speed, aggregate throughput and time to first token for self-hosted models

An LLM’s tokens per second is capped by memory bandwidth: every generated token reads the active weights and the conversation’s KV cache once. Usable bandwidth divided by those bytes gives the ceiling, and prompt processing is capped by compute instead. At the defaults below, Qwen3-32B at FP8 on one H100 SXM 80 GB tops out near 60 tokens/s for one user, with the first token after about 136 ms for a 2,048-token prompt. Every figure is an estimate and an upper bound.

Estimate tokens per second for your model and GPU

Choose a model, its precision, the prompt and response lengths, how many requests run at once and the hardware. Results recalculate in your browser as you change any field; nothing is sent anywhere.

Each change updates the estimate. A short summary is announced when you stop typing.

Model

Dense, grouped-query attention. 64 layers, 8 KV heads of dimension 128. Qwen3-32B config.json

Precision
Workload

Sequences decoding together. For a team, 10–20% of named users is a common peak.

Hardware

80 GB HBM3, 3.35 TB/s. Dense TFLOPS: FP16 989.5, FP8 1,979 (half of the 1,979 and 3,958 printed with sparsity). H100 specs

GPUs in one server use tensor parallelism; desk-side systems link by layers.

Assumptions
Assumptions

0.60: Databricks measured 60% on two H100s at batch size 1.

0.50: our assumption, below the 76% Google reported for large-batch prefill.

0.90: our assumption for the share of summed bandwidth kept under tensor parallelism.

Auto runs FP8 checkpoints on FP8 Tensor Cores, NVFP4 on FP4, and dequantized formats at 16-bit.

Estimated speed (upper bound)

60 tokens/s per user while generating

Memory-bandwidth bound Memory-bandwidth bound: each step waits on streaming the weights and KV cache, not on the maths. The matrix maths would overtake the weight read at about 247 concurrent sequences, but memory holds only 58 of this length, so memory caps the batch first. Reading the KV cache takes 2% of each step by the last token.

Fits: needs 37.1 GB of 80 GB. Size the memory in the VRAM calculator

Aggregate decode throughput
60 tokens/s 1 concurrent request
Time to first token
136 ms one request on an idle server, FP8 maths
Time to generate the response
8.6 s 512 tokens: first token plus 8.5 s of decoding
Prefill if every request arrives at once
136 ms
Weights read per decode step
32.8 GB FP8 or INT8 (8 bits per weight)
KV cache read per request, last token
0.7 GB
Speed at the first and last token
60 tokens/s → 60 tokens/s

Speed at each concurrency

One weight read serves the whole batch, so total throughput climbs with concurrency until the matrix maths outlasts that read. Rows past the memory limit show what the speed would be if the KV cache fitted.

Per-user and aggregate decode speed by number of concurrent requests
Concurrent requestsPer userAggregateLimitMemory
1 60 tokens/s 60 tokens/s Bandwidth Fits
2 59 tokens/s 118 tokens/s Bandwidth Fits
4 57 tokens/s 229 tokens/s Bandwidth Fits
8 53 tokens/s 428 tokens/s Bandwidth Fits
16 47 tokens/s 758 tokens/s Bandwidth Fits
32 39 tokens/s 1,235 tokens/s Bandwidth Fits
64 28 tokens/s 1,801 tokens/s Bandwidth Does not fit
128 18 tokens/s 2,337 tokens/s Bandwidth Does not fit
256 11 tokens/s 2,727 tokens/s Compute Does not fit

How the calculator works

Serving a request has two phases with different bottlenecks. Prefill pushes the whole prompt through the model in one parallel pass and is limited by arithmetic; decode then produces one token per step and is limited by how fast the GPU can stream bytes from memory. Databricks and NVIDIA both describe the split this way, and kipply’s inference arithmetic supplies the step formulas used here.

  1. Bytes per decode step. Active parameters × bits per weight ÷ 8, plus each request’s KV cache at its current length. The cache uses the same per-layer formula as the VRAM calculator: 2 × layers × KV heads × head dimension × bytes per value per token, with sliding-window layers capped at their window.
  2. Effective bandwidth. Vendor bandwidth × GPUs × multi-GPU efficiency × MBU. Model bandwidth utilisation is the share of peak bandwidth a real engine achieves.
  3. Decode step time. The larger of weight bytes ÷ effective bandwidth and 2 × active parameters × batch ÷ effective compute, plus the batch’s KV bytes ÷ effective bandwidth. Attention reads every request’s cache separately, so that term never shrinks with batching.
  4. Per-user speed and aggregate. Tokens 2 to N each take one step as the context grows from the prompt length to prompt + N. Per-user speed is 1 ÷ the average step time; aggregate throughput multiplies it by the number of concurrent requests.
  5. Crossover. Weights stream once per step whatever the batch, while the maths grows with it. For a dense model the two meet at a batch of bytes per weight × dense FLOPS × MFU ÷ (2 × bandwidth × MBU). For FP8 weights on an H200 that is 1 × 989.5 TFLOPS ÷ (2 × 4.8 TB/s × 0.6), about 172 requests; kipply’s A100 example puts the same ratio at 208 for 16-bit weights at full utilisation.
  6. Time to first token. The larger of 2 × active parameters × prompt tokens ÷ effective compute and one read of the weights, plus writing the prompt’s KV cache. Effective compute is dense TFLOPS × GPUs × multi-GPU efficiency × MFU.
  7. Mixture of experts. One token reads the shared layers plus its chosen experts, which adds up to the active parameter count. A batch touches more experts: with E experts, k chosen per token and b tokens, the expected share touched is 1 − (1 − k/E)b, so weight reads climb from the active count toward the total.
  8. Memory fit. Weights plus the KV cache for every request at its full length, divided by a usable share of 0.90 (0.80 on unified-memory systems), checked against the GPUs’ memory. When it does not fit, the result says so and points to the VRAM calculator.

A check by hand: Llama 3.3 70B at FP8 reads 70.6 GB of weights per token. An H200’s 4.8 TB/s at 60% utilisation is 2.88 TB/s, so a token takes about 24.5 ms before the KV cache is counted, about 41 tokens per second. The cache at a 2,300-token context adds roughly 0.26 ms, which brings the estimate to the 40 tokens/s in the worked examples.

Default assumptions and where they come from

SettingDefaultBasis
Bandwidth utilisation (MBU)60%Databricks reported 60% MBU on two H100-80GB GPUs at batch size 1, and 55% on four A100-40GB GPUs (Databricks, October 2023).
Compute utilisation (MFU)50%Our assumption. Google reached 76% MFU on large-batch prompt processing for PaLM 540B (Pope et al., 2022); single requests and GPU servers usually run the hardware less fully, so the default sits lower.
Multi-GPU efficiency90%Our assumption. Tensor parallelism adds an all-reduce to every layer (kipply), and Databricks saw MBU fall as small models were spread over more GPUs.
Dense computeVendor figureSparse ratings are halved, because 2:4 structured sparsity doubles the rated Tensor Core throughput (NVIDIA). Weight-only formats such as AWQ, GPTQ, GGUF and MXFP4 are assumed to multiply at 16-bit.
Usable memory for the fit check90%, or 80% on unified memoryThe same reserve the VRAM calculator holds back for activations, the runtime and, on desk-side systems, the operating system.

Worked examples

Four cases at a 2,048-token prompt, a 512-token response and a 16-bit KV cache, computed when this page was built by the same module the calculator runs. The build script recomputes each one by hand and fails if they disagree.

Estimates and upper bounds at the default utilisation settings (MBU 0.60, MFU 0.50)
Scenario Hardware Per user Aggregate First token Full response Limit
Llama 3.3 70B, FP8, one user 1 × H200 SXM 40 tokens/s 40 tokens/s 292 ms 13 s Bandwidth
Qwen3-8B, BF16, one user 1 × RTX 5090 64 tokens/s 64 tokens/s 321 ms 8.3 s Bandwidth
gpt-oss-120b as published (MXFP4 experts), one user 1 × DGX Spark 33 tokens/s 33 tokens/s n/a n/a Bandwidth (compute unknown)
Llama 3.3 70B, FP8, 32 concurrent users 1 × H200 SXM 30 tokens/s 973 tokens/s 292 ms 17 s Bandwidth
  • gpt-oss-120b keeps attention, embeddings and the output layer in BF16 and only its experts in MXFP4, so one token reads 4.9 GB rather than the 2.7 GB a uniform 4.25-bit average would suggest.
  • DGX Spark’s first-token time is n/a because NVIDIA publishes only an FP4 rating with sparsity for it; choose NVFP4 weights, or a custom GPU with a measured figure, to get one.
  • At 32 concurrent users the H200 still waits on memory: the crossover is about 172 requests, but its memory holds only 67 requests of this length, and by the last token the KV cache takes 28% of each step.

Memory sizing comes before speed. The GPU requirements walkthrough turns a headcount into concurrent requests and GPUs, which is the concurrency to enter here.

GPU bandwidth and compute

Bandwidth decides decode speed and dense Tensor Core throughput decides prefill. All figures below were read from each vendor’s spec page or architecture document (verified October 2026) and stored without sparsity.

Accelerator presets, dense TFLOPS and sources
Accelerator Memory Bandwidth Dense TFLOPS Source
NVIDIA L40S 48 GB GDDR6 864 GB/s FP16 362.05, FP8 733 (dense; NVIDIA prints dense | sparse) L40S specs
NVIDIA GeForce RTX 5090 32 GB GDDR7 1,792 GB/s FP16 209.5, FP8 419, FP4 1,676 (dense, FP32 accumulate) RTX 5090 specs
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB GDDR7 1,792 GB/s FP16 503.8, FP8 1,007.6, FP4 2,015.2 (dense, FP32 accumulate) RTX PRO 6000 Workstation specs
NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB GDDR7 1,597 GB/s FP16 500, FP8 1,000, FP4 2,000 (half of the 1, 2 and 4 PFLOPS printed) RTX PRO 6000 Server specs
NVIDIA H100 SXM 80 GB HBM3 3.35 TB/s FP16 989.5, FP8 1,979 (half of the 1,979 and 3,958 printed with sparsity) H100 specs
NVIDIA H100 NVL (PCIe) 94 GB HBM3 3.9 TB/s FP16 835.5, FP8 1,670.5 (half of the 1,671 and 3,341 printed with sparsity) H100 NVL specs
NVIDIA H200 SXM 141 GB HBM3e 4.8 TB/s FP16 989.5, FP8 1,979 (half of the figures printed with sparsity) H200 specs
NVIDIA H200 NVL (PCIe) 141 GB HBM3e 4.8 TB/s FP16 835.5, FP8 1,670.5 (half of the figures printed with sparsity) H200 NVL specs
NVIDIA B200 (one GPU of an 8-GPU system) 180 GB HBM3e 8 TB/s FP16 2,250, FP8 4,500, FP4 9,000 (dense, eight-GPU totals ÷ 8) DGX B200 specs
AMD Instinct MI300X 192 GB HBM3 5.3 TB/s FP16 1.3 PFLOPS, FP8 2.61 PFLOPS (dense, as printed) MI300X specs
AMD Instinct MI325X 256 GB HBM3E 6 TB/s FP16 1.3 PFLOPS, FP8 2.61 PFLOPS (dense, as printed) MI325X specs
AMD Instinct MI355X 288 GB HBM3E 8 TB/s FP16 2.5 PF, FP8 5.0 PF, MXFP4 10 PF (dense) MI350 Series reference
NVIDIA DGX Spark (128 GB unified) 128 GB LPDDR5x, unified 273 GB/s FP4 500 (half of “1 PFLOP at FP4 with sparsity”); no FP16 or FP8 figure published DGX Spark hardware overview
Apple Mac Studio, M5 Max (40-core GPU) with 128 GB 128 GB unified 614 GB/s Apple publishes no GPU TFLOPS figure Mac Studio specs
Apple Mac Studio, M5 Ultra with 256 GB 256 GB unified 1.2 TB/s Apple publishes no GPU TFLOPS figure Mac Studio specs
Apple Mac Studio, M5 Ultra with 512 GB 512 GB unified 1.2 TB/s Apple publishes no GPU TFLOPS figure Mac Studio specs
AMD Ryzen AI Max+ 395 system with 128 GB 128 GB LPDDR5x-8000, unified 256 GB/s AMD publishes no GPU TFLOPS figure for the Radeon 8060S Ryzen AI Halo specs

Against the same model, the decode ceiling moves with bandwidth. Llama 3.3 70B at FP8 comes out near 2.3 tokens/s on a 273 GB/s DGX Spark and 10 tokens/s on a 1.2 TB/s M5 Ultra Mac Studio. The desk-side system comparison covers the software and clustering differences, and the server buying guide covers power, cooling and interconnects for rack GPUs.

What the estimate leaves out

The calculator answers whether a model and GPU can plausibly reach a target speed. It cannot promise that a given engine will. These effects lower real numbers or change them in ways the arithmetic does not see:

  • Engine overhead. Kernel launches, sampling, tokenization, scheduling and the HTTP layer cost time on every step, and the gap is widest on small models.
  • Attention FLOPs. Prefill counts 2 × parameters per token. The attention term grows with the square of the prompt, so very long prompts take noticeably longer than shown.
  • Batching policy. Real servers mix prefill and decode in the same steps, and continuous batching means requests rarely start together. Queueing adds to time to first token under load.
  • Speculative decoding and prefix caching. Draft models, multi-token prediction and shared prefixes can beat the ceiling on some workloads. The calculator assumes none of them.
  • Routing skew. The expert-coverage term assumes tokens spread evenly over experts, the case that reads the most weights. Skewed routers read fewer.
  • Desk-side pipelines. Linked systems are modelled as one unit’s bandwidth. With many users, pipeline stages can overlap and raise total throughput.

Confirm a shortlist with a load test. The serving engine comparison shows how to step concurrency with vllm bench serve and read time to first token, inter-token latency and total throughput from the run.

How VDF AI fits

VDF AI is the platform layer that runs on GPUs you have sized, on-premises, in a private cloud or air-gapped. Its model router registers Ollama and other on-prem deployments next to any permitted cloud models, and its routing policies can set latency limits on time to first token and tokens per second.

An air-gap mode restricts routing to local models only.

Tokens per second questions

How do I calculate tokens per second for an LLM?

Divide the memory bandwidth you can actually use by the bytes each generated token has to read. Those bytes are the active weights plus the conversation’s KV cache. Llama 3.3 70B at FP8 reads about 70.6 GB per token; an H200 moves 4.8 TB/s, and at 60% bandwidth utilisation that is 2.88 TB/s, so one user tops out near 40 tokens/s. Prefill is different: it is limited by compute, at 2 × parameters × prompt tokens FLOPs. Treat both results as upper bounds.

How many tokens per second can an H100 generate?

It depends on the model and on how many requests share the GPU. Qwen3-32B at FP8 on one H100 SXM tops out near 60 tokens/s for a single user, because every token streams 32.8 GB of weights. With 32 concurrent users each one gets about 39 tokens/s, but the GPU produces about 1,235 tokens/s in total, since one weight read serves the whole batch. A smaller or 4-bit model goes faster; a 70B model at FP8 barely fits in 80 GB.

Why is prompt processing so much faster than generation?

Prefill handles every prompt token in one parallel pass, so the GPU’s arithmetic units stay busy, while generation produces one token per step and spends most of each step waiting on memory. On one H100 with Qwen3-32B at FP8, the calculator puts a 2,048-token prompt at about 136 ms, roughly 15,000 prompt tokens per second, against 60 tokens/s of output. Long prompts in retrieval and agent workloads therefore show up as time to first token, not as slower generation.

Do more GPUs increase tokens per second?

Inside one server, yes: tensor parallelism splits every layer, so each GPU reads only its share of the weights. The calculator adds the cards’ bandwidth and keeps 90% of the sum to allow for the all-reduce traffic. Llama 3.3 70B at FP8 goes from about 28 tokens/s on one H100 SXM to 51 tokens/s on two. Linked desk-side systems such as DGX Spark or Mac Studio split the model by layers instead, which adds memory but does not speed up a single conversation.

Are mixture-of-experts models faster at the same size?

For one user, much faster. gpt-oss-120b keeps 117 billion parameters loaded but each token reads only about 4.9 GB as published, so one DGX Spark can reach about 33 tokens/s. The advantage shrinks with concurrency: different users’ tokens pick different experts, and 64 concurrent sequences on an H100 touch about 56.2 GB of weights per step. Memory still has to hold every expert, so size capacity on the total and speed on the active share.

Why are my measured tokens per second lower than the estimate?

The estimate is a ceiling. Real servers lose time to kernel launches, sampling, scheduling, network hops and queueing, and the default 60% bandwidth utilisation may be optimistic for your engine, especially on small models or many GPUs. Quantized formats that need dequantizing add overhead, and long prompts add attention work the estimate leaves out. Measure with your own prompts and concurrency, for example with vllm bench serve, then set the utilisation fields to match what you observe.