AI Infrastructure

RTX PRO 6000 vs H100 vs H200 vs B200 for LLM Inference (2026)

NVIDIA's RTX PRO 6000 Blackwell, H100, H200 and B200 compared from NVIDIA's own datasheets: memory, bandwidth, FP8 and FP4 Tensor Core throughput, NVLink or PCIe, power and form factor, plus sessions-per-GPU arithmetic and bandwidth-bound speed ceilings for 32B and 70B models.

RTX PRO 6000 vs H100 comes down to capacity against bandwidth. NVIDIA's RTX PRO 6000 Blackwell puts 96 GB at 1.8 TB/s on a PCIe card. The H100 SXM has 80 GB at 3.35 TB/s with NVLink, the H200 raises that to 141 GB at 4.8 TB/s, and a B200 reaches 180 GB at 7.7 TB/s and adds FP4.

These four GPUs cover most of what enterprises buy in 2026 to serve open-weight models on their own hardware. The RTX PRO 6000 Blackwell is a professional card that fits ordinary PCIe workstations and servers. H100 and H200 are Hopper-generation data-centre GPUs, sold as SXM modules on four- or eight-GPU boards and as dual-slot PCIe NVL cards. B200 is a Blackwell-generation data-centre GPU that ships on eight-GPU HGX and DGX systems.

Every specification here comes from NVIDIA’s product pages and datasheets, read in October 2026. The post contains no tokens-per-second measurements of its own. Where it discusses speed, it either derives a bandwidth-bound ceiling with the formula shown or quotes an NVIDIA claim together with its test conditions.

Quick picks

Your priorityPickWhy
A 70B model or gpt-oss-120b in a standard air-cooled serverRTX PRO 6000 Blackwell (Server Edition or Max-Q)96 GB on one PCIe card, FP8 and FP4, up to four MIG instances
Fast streams from servers you already ownH100 SXM or H100 NVL3.35 to 3.9 TB/s and NVLink
Many concurrent sessions on a 70B-class modelH200141 GB at 4.8 TB/s holds the weights and a deep KV cache
The largest models and FP4 at data-centre scaleB200, or B300 for more memory180 GB or 270 GB at 7.7 TB/s, NVLink at 1.8 TB/s
One developer running models up to about 32B at 4-bitRTX 5090The same 1,792 GB/s as the RTX PRO 6000, with 32 GB

Specifications side by side

Figures per GPU from NVIDIA’s product pages, datasheets and RTX Blackwell architecture whitepaper (verified October 2026). Where a page lists SXM and PCIe versions, both are shown as SXM / NVL.

RTX 5090RTX PRO 6000 BlackwellH100 SXM / NVLH200 SXM / NVLB200 (HGX)B300 (HGX)
Memory32 GB GDDR796 GB GDDR7 with ECC80 GB / 94 GB141 GB HBM3e180 GB HBM3E270 GB HBM3E
Bandwidth1,792 GB/s1,792 GB/s; 1,597 GB/s on the Server Edition3.35 / 3.9 TB/s4.8 TB/s7.7 TB/s7.7 TB/s
FP8 Tensor Core, sparse1,676 TFLOPS (FP16 accumulate)2 PFLOPS (Server Edition page)3,958 / 3,341 TFLOPS3,958 / 3,341 TFLOPS9 PFLOPS9 PFLOPS
FP4 Tensor Core, sparse / dense3,352 / 1,676 TFLOPS4,000 TOPS sparseNot listedNot listed18 / 9 PFLOPS18 / 14 PFLOPS
GPU-to-GPU linkPCIe 5.0PCIe 5.0 x16; no NVLink listedNVLink 900 GB/s / NVLink 600 GB/sNVLink 900 GB/s / NVLink bridge 900 GB/sNVLink 5, 1.8 TB/sNVLink 5, 1.8 TB/s
Maximum power575 W600 W; Max-Q 300 W; Server up to 600 W, configurableUp to 700 W / 350–400 WUp to 700 W / up to 600 WUp to 1,000 WUp to 1,100 W
Form factorDual-slot PCIeDual-slot PCIe; single-slot liquid-cooled Server optionSXM module / dual-slot PCIeSXM module / dual-slot PCIeSXM on an eight-GPU boardSXM on an eight-GPU board

Notes behind the table:

  • Sparse figures assume structured sparsity. NVIDIA’s Blackwell datasheets state that dense FP8 throughput is half the sparse number shown. The RTX PRO 6000 Server Edition page lists 4 PFLOPS of FP4 and 2 PFLOPS of FP8 without a footnote; the Workstation Edition datasheet marks its matching 4,000 TOPS figure as using sparsity.
  • Partitioning. The RTX PRO 6000 supports up to four MIG instances (four of 24 GB on the Workstation Edition). H100 and H200 support up to seven, at 10 GB and 18 GB each on the SXM versions, and NVIDIA’s Blackwell datasheet also lists seven for the B200.
  • System totals differ slightly. A DGX B200 lists 1,440 GB and 64 TB/s across eight GPUs, so its bandwidth works out a little above the 7.7 TB/s per GPU in the HGX B200 datasheet.
  • No list prices for the data-centre parts. NVIDIA’s H100, H200 and DGX B200 pages show no prices, so budgets start from partner quotes. The RTX PRO 6000 is the exception: NVIDIA’s US marketplace listed it at $16,000 (list, verified October 2026).

Bandwidth sets the speed ceiling

Generating a token means reading every active weight from memory once, plus that conversation’s KV cache. One stream therefore cannot decode faster than memory bandwidth ÷ bytes read per token. For a dense model at a short context, the bytes read are close to the size of the weights.

Ceilings for one stream, in tokens per second, from the bandwidth figures above and published parameter counts. Weights use the same arithmetic as our VRAM calculator: 8 bits per weight for FP8 and about 4.25 for AWQ or GPTQ 4-bit.

GPULlama 3.3 70B, FP8 (70.6 GB)Llama 3.3 70B, 4-bit (37.5 GB)Qwen3-32B, FP8 (32.8 GB)Qwen3-32B, 4-bit (17.4 GB)
RTX 5090Does not fitDoes not fitDoes not fit≈ 103
RTX PRO 6000 Workstation≈ 25≈ 48≈ 55≈ 103
H100 SXM≈ 47, with almost no room for KV cache≈ 89≈ 102≈ 192
H200 SXM≈ 68≈ 128≈ 147≈ 276
B200≈ 109≈ 205≈ 235≈ 442

Three things keep real numbers below these ceilings. Kernels never reach the full rated bandwidth, the KV cache adds bytes to every step as a conversation grows, and serving engines spend time scheduling. The useful signal is the ratio between GPUs: for the same model, a B200 can stream about four times as fast as an RTX PRO 6000, and an H200 about 1.4 times as fast as an H100 SXM.

Batching shifts the picture. A serving engine decodes many sequences in one step, and the weights are read once per step for all of them. With five 8,192-token sessions of Llama 3.3 70B at FP8 on an RTX PRO 6000, a full step reads about 70.6 GB of weights plus up to 13.4 GB of cache. The aggregate ceiling is then about 5 × 1,792 ÷ 84 ≈ 107 tokens per second, shared by the five users. At larger batches the work becomes limited by compute, which is where the FP8 and FP4 rows of the table start to count.

Mixture-of-experts models read only their active experts for each token. gpt-oss-120b keeps 117 billion parameters in memory but activates 5.1 billion per token, so its decode ceiling sits far above a dense model of the same footprint.

How many sessions fit on each GPU

Speed is one half of the sizing question; the other is how many conversations stay resident at once. The method matches our GPU sizing guide:

  1. KV cache per token = 2 × layers × KV heads × head dimension × bytes per value. Llama 3.3 70B has 80 layers with 8 KV heads of dimension 128, so a 16-bit cache costs 327,680 bytes per token, or 2.68 GB for an 8,192-token session. Qwen3-32B’s 64 layers cost 262,144 bytes per token, or 2.15 GB per session.
  2. Sessions that fit = (GPU memory × 0.9 − weights) ÷ KV cache per session, which leaves a tenth of memory for activations and the runtime.
  3. Users ≈ sessions ÷ 0.1 to 0.2, since 10 to 20 percent of named users tend to be active at the busiest moment.

8,192-token sessions that fit, with a 16-bit KV cache / an FP8 KV cache:

GPULlama 3.3 70B, FP8Llama 3.3 70B, 4-bitQwen3-32B, FP8gpt-oss-120b as shipped (65.2 GB)
RTX 5090, 32 GBDoes not fitDoes not fitDoes not fitDoes not fit
RTX PRO 6000, 96 GB5 / 1118 / 3624 / 4968 / 137
H100 SXM, 80 GB0 / 112 / 2518 / 3622 / 44
H100 NVL, 94 GB5 / 1017 / 3524 / 4863 / 126
H200, 141 GB20 / 4133 / 6643 / 87201 / 402
B200, 180 GB34 / 6846 / 9260 / 120315 / 630
B300, 270 GB64 / 12876 / 15397 / 195579 / 1,159

The gpt-oss-120b column uses the checkpoint as published, with MXFP4 experts. Only half of its 36 layers keep a full-length KV cache; the other half keep a 128-token sliding window. Context length moves every cell: at 32,768 tokens, Llama 3.3 70B at FP8 fits 1 session on an RTX PRO 6000, 5 on an H200 and 8 on a B200. To try other models and contexts, use the calculator.

Two patterns stand out. At FP8, a 70B model leaves an 80 GB H100 SXM with no room for conversations, which is why so many 70B deployments run 4-bit weights or move to H200. And the RTX PRO 6000 sits between H100 and H200 on capacity, so it serves gpt-oss-120b to dozens of sessions even though each stream is slower.

RTX PRO 6000 vs RTX 5090

The two cards share a memory bandwidth figure, 1,792 GB/s, so the question is almost entirely about capacity and operations:

  • Memory. 96 GB with error correction against 32 GB without it. On the 5090, a 32B model at 4-bit holds four or five 8,192-token sessions; the RTX PRO 6000 holds about thirty, or a 70B model.
  • Compute. The RTX PRO 6000 lists 24,064 CUDA cores and 4,000 sparse FP4 TOPS. The RTX 5090 lists 21,760 CUDA cores and 3,352 sparse FP4 TFLOPS.
  • Sharing. MIG on the RTX PRO 6000 splits the card into up to four isolated 24 GB instances, which lets several teams or small models share one GPU. The 5090 has no MIG.
  • Power and form. 600 W and 575 W respectively. The 300 W Max-Q version keeps the full 96 GB and bandwidth for multi-GPU workstations, and the Server Edition comes as a dual-slot air-cooled or single-slot liquid-cooled card at 1,597 GB/s.
  • Price. NVIDIA’s US list price for the RTX 5090 Founders Edition is $1,999, and its marketplace listed the RTX PRO 6000 at $16,000 (list, verified October 2026). Street prices for the 5090 ran far above list in 2026: Hardware Busters reported Micro Center asking around $4,200 in September.

If one person runs models that fit in 32 GB, the 5090 delivers the same per-stream ceiling for a fraction of the price. Shared serving, 70B-class models and anything that needs ECC or isolation points to the RTX PRO 6000. Our GPU buyer’s guide for local models compares the consumer cards in more detail.

H200 vs H100

Same architecture, different memory. On the SXM versions NVIDIA lists identical Tensor Core peaks, 3,958 sparse FP8 TFLOPS, and the same 700 W ceiling. What changes:

  • Capacity: 141 GB against 80 GB, which NVIDIA describes as nearly double.
  • Bandwidth: 4.8 TB/s against 3.35 TB/s, about 1.4 times higher, so the per-stream ceiling rises by the same factor.
  • MIG slices: seven instances of 18 GB on H200 against seven of 10 GB on H100.
  • PCIe versions: H200 NVL keeps 141 GB and 4.8 TB/s at up to 600 W; H100 NVL has 94 GB and 3.9 TB/s at 350–400 W.

NVIDIA’s H200 page claims 1.9 times the Llama 2 70B inference throughput of an H100. The footnote explains much of it: the test ran at batch size 8 on the H100 and batch size 32 on the H200, with 2,000 input and 128 output tokens. More memory means bigger batches, and bigger batches mean more tokens per GPU. That is the case for H200 in one sentence. If your H100s serve 70B-class models with only a handful of sessions each, H200 is the like-for-like upgrade within the same HGX server design.

H200 vs B200, and where B300 fits

The B200 is a larger step. Per GPU, from NVIDIA’s Blackwell datasheet:

  • Memory and bandwidth: 180 GB at 7.7 TB/s, against 141 GB at 4.8 TB/s.
  • Tensor Cores: 9 PFLOPS of sparse FP8, against 3.958, and FP4 at 18 PFLOPS sparse or 9 dense, which Hopper does not list.
  • NVLink: fifth generation at 1.8 TB/s per GPU, twice the H200’s 900 GB/s.
  • Power: configurable up to 1,000 W per GPU. A DGX B200 is a 10 RU system rated at about 14.3 kW.

NVIDIA’s DGX B200 page claims 15 times the inference performance of DGX H100, as a projected per-GPU figure for real-time serving with 50 ms between tokens, 32,768 input tokens and 1,028 output tokens. Treat it as a ceiling for that workload, not a general multiplier.

The B300, built on Blackwell Ultra, keeps the 7.7 TB/s and the 9 sparse FP8 PFLOPS but raises memory to 270 GB and dense FP4 to 14 PFLOPS, at up to 1,100 W per GPU. A DGX B300 is rated at about 14 kW and moves to 800 Gb/s ConnectX-8 networking. For memory-bound serving of very large models, B300’s extra 90 GB per GPU matters more than its compute. NVIDIA’s HGX page already lists the next generation, Rubin, at 288 GB of HBM4 and 22 TB/s per GPU.

Decision table: PCIe workstation card or SXM data-centre GPU

SituationBetter fitReason
One model up to about 80 GB, a few dozen named usersRTX PRO 6000 in a PCIe server96 GB per card without SXM boards or liquid cooling
Several small models or teams that must stay isolatedRTX PRO 6000 with MIGUp to four instances with their own memory
Racks limited to air cooling and modest power per serverRTX PRO 6000 Max-Q, H100 NVL or H200 NVL300 W, 350–400 W or up to 600 W per card
A model larger than one GPU, split with tensor parallelismH200 or B200 on HGXNVLink at 900 GB/s or 1.8 TB/s; NVIDIA lists no NVLink for the RTX PRO 6000
Hundreds of concurrent or long-context sessionsH200, B200 or B300141 to 270 GB per GPU and 4.8 to 7.7 TB/s
FP4 serving at data-centre scaleB200 or B30018 PFLOPS of sparse FP4 per GPU
A first pilot on a single serverRTX PRO 6000 or H100 NVLStandard PCIe hosts, smaller power envelope

The pattern behind the table: PCIe cards win when a model fits on one card and the facility is built for ordinary servers. SXM GPUs win when the model, the context or the user count outgrows one card, because NVLink keeps multi-GPU serving efficient. Facilities often decide before models do. Eight 1,000 W GPUs plus their host need data-centre power and cooling, which the on-prem server guide covers, and the buy, lease or colocate analysis helps when your own rooms cannot take that load.

How VDF AI fits

VDF AI’s self-hosted package runs as containers, and its documented sizing for the platform itself lists CPU, memory and storage: 4 to 8 vCPUs and 16 to 32 GB of memory for a single-host pilot. The GPUs compared here serve the models; the platform does not consume them.

VDF AI Router keeps applications independent of that hardware choice. Ollama and custom on-premises deployments sit in its model registry next to any cloud models your policy permits, and the router prefers local models when they are available. An RTX PRO 6000 pilot endpoint can later be joined or replaced by H200 or B200 endpoints without changing the applications that call the router. In air-gap mode, routing is restricted to local models only.

Sources

Verified 6 October 2026.


Sizing GPUs for a private model service? Start with the reference architecture or book a demo.

Frequently asked questions

Is the RTX PRO 6000 better than the H100 for LLM inference?

It depends on which limit you reach first. The RTX PRO 6000 Blackwell has more memory, 96 GB against 80 GB on an H100 SXM, and adds FP4, so it can hold a 70B model at FP8 with a few sessions to spare. The H100 SXM has nearly twice the memory bandwidth, 3.35 TB/s against 1.79 TB/s, plus 900 GB/s NVLink, so each user gets tokens faster and multi-GPU models scale better. Pick the PCIe card for capacity in ordinary servers and the H100 for speed per stream.

What is the difference between H200 and H100?

Both are Hopper GPUs, and NVIDIA lists the same FP8 Tensor Core peak for the SXM versions: 3,958 teraFLOPS with sparsity. The H200 changes the memory, with 141 GB of HBM3e at 4.8 TB/s against 80 GB at 3.35 TB/s. One H200 can hold a 70B model at FP8 plus about twenty 8,000-token sessions, while an H100 SXM holds the weights and almost nothing else. NVIDIA reports 1.9 times the Llama 2 70B throughput, measured with a four times larger batch on the H200.

Is the B200 worth it over the H200?

For large or busy deployments, usually yes. A B200 has 180 GB at 7.7 TB/s, adds FP4 Tensor Cores and doubles NVLink bandwidth to 1.8 TB/s per GPU, and NVIDIA lists 9 petaFLOPS of sparse FP8 per GPU against 3.958 for an H200. The trade-offs are power, up to 1,000 W per GPU and about 14.3 kW for a DGX B200, and packaging, since B200s ship on eight-GPU boards. For one model on a few GPUs in standard servers, an H200 NVL can be enough.

Can an RTX 5090 replace an RTX PRO 6000 for local inference?

Only for models that fit in 32 GB. Both cards list 1,792 GB/s of memory bandwidth, so a model that fits on the RTX 5090 has the same bandwidth ceiling on either card. The RTX PRO 6000 has three times the memory, error-correcting memory and Multi-Instance GPU partitioning, which matter for 70B-class models, long contexts and shared servers. A 32B model at 4-bit leaves room for about four or five 8,000-token sessions on a 5090 and about thirty on an RTX PRO 6000.

How many users can one H200 or RTX PRO 6000 serve?

Count concurrent sessions rather than users, and assume 10 to 20 percent of named users are active at once. With Llama 3.3 70B at FP8 and 8,000-token sessions, the memory arithmetic allows about 5 sessions on an RTX PRO 6000 and about 20 on an H200, roughly doubling with an FP8 KV cache. With 4-bit weights the RTX PRO 6000 holds about 18. Speed per user falls as sessions share bandwidth, so load-test your own prompts before buying.

Filed under
AI infrastructureGPU capacity planningLLM inferencelocal AI hardwareon-premises AImodel serving
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading