RTX PRO 6000 vs H100 comes down to capacity against bandwidth. NVIDIA's RTX PRO 6000 Blackwell puts 96 GB at 1.8 TB/s on a PCIe card. The H100 SXM has 80 GB at 3.35 TB/s with NVLink, the H200 raises that to 141 GB at 4.8 TB/s, and a B200 reaches 180 GB at 7.7 TB/s and adds FP4.
These four GPUs cover most of what enterprises buy in 2026 to serve open-weight models on their own hardware. The RTX PRO 6000 Blackwell is a professional card that fits ordinary PCIe workstations and servers. H100 and H200 are Hopper-generation data-centre GPUs, sold as SXM modules on four- or eight-GPU boards and as dual-slot PCIe NVL cards. B200 is a Blackwell-generation data-centre GPU that ships on eight-GPU HGX and DGX systems.
Every specification here comes from NVIDIA’s product pages and datasheets, read in October 2026. The post contains no tokens-per-second measurements of its own. Where it discusses speed, it either derives a bandwidth-bound ceiling with the formula shown or quotes an NVIDIA claim together with its test conditions.
Quick picks
| Your priority | Pick | Why |
|---|---|---|
| A 70B model or gpt-oss-120b in a standard air-cooled server | RTX PRO 6000 Blackwell (Server Edition or Max-Q) | 96 GB on one PCIe card, FP8 and FP4, up to four MIG instances |
| Fast streams from servers you already own | H100 SXM or H100 NVL | 3.35 to 3.9 TB/s and NVLink |
| Many concurrent sessions on a 70B-class model | H200 | 141 GB at 4.8 TB/s holds the weights and a deep KV cache |
| The largest models and FP4 at data-centre scale | B200, or B300 for more memory | 180 GB or 270 GB at 7.7 TB/s, NVLink at 1.8 TB/s |
| One developer running models up to about 32B at 4-bit | RTX 5090 | The same 1,792 GB/s as the RTX PRO 6000, with 32 GB |
Specifications side by side
Figures per GPU from NVIDIA’s product pages, datasheets and RTX Blackwell architecture whitepaper (verified October 2026). Where a page lists SXM and PCIe versions, both are shown as SXM / NVL.
| RTX 5090 | RTX PRO 6000 Blackwell | H100 SXM / NVL | H200 SXM / NVL | B200 (HGX) | B300 (HGX) | |
|---|---|---|---|---|---|---|
| Memory | 32 GB GDDR7 | 96 GB GDDR7 with ECC | 80 GB / 94 GB | 141 GB HBM3e | 180 GB HBM3E | 270 GB HBM3E |
| Bandwidth | 1,792 GB/s | 1,792 GB/s; 1,597 GB/s on the Server Edition | 3.35 / 3.9 TB/s | 4.8 TB/s | 7.7 TB/s | 7.7 TB/s |
| FP8 Tensor Core, sparse | 1,676 TFLOPS (FP16 accumulate) | 2 PFLOPS (Server Edition page) | 3,958 / 3,341 TFLOPS | 3,958 / 3,341 TFLOPS | 9 PFLOPS | 9 PFLOPS |
| FP4 Tensor Core, sparse / dense | 3,352 / 1,676 TFLOPS | 4,000 TOPS sparse | Not listed | Not listed | 18 / 9 PFLOPS | 18 / 14 PFLOPS |
| GPU-to-GPU link | PCIe 5.0 | PCIe 5.0 x16; no NVLink listed | NVLink 900 GB/s / NVLink 600 GB/s | NVLink 900 GB/s / NVLink bridge 900 GB/s | NVLink 5, 1.8 TB/s | NVLink 5, 1.8 TB/s |
| Maximum power | 575 W | 600 W; Max-Q 300 W; Server up to 600 W, configurable | Up to 700 W / 350–400 W | Up to 700 W / up to 600 W | Up to 1,000 W | Up to 1,100 W |
| Form factor | Dual-slot PCIe | Dual-slot PCIe; single-slot liquid-cooled Server option | SXM module / dual-slot PCIe | SXM module / dual-slot PCIe | SXM on an eight-GPU board | SXM on an eight-GPU board |
Notes behind the table:
- Sparse figures assume structured sparsity. NVIDIA’s Blackwell datasheets state that dense FP8 throughput is half the sparse number shown. The RTX PRO 6000 Server Edition page lists 4 PFLOPS of FP4 and 2 PFLOPS of FP8 without a footnote; the Workstation Edition datasheet marks its matching 4,000 TOPS figure as using sparsity.
- Partitioning. The RTX PRO 6000 supports up to four MIG instances (four of 24 GB on the Workstation Edition). H100 and H200 support up to seven, at 10 GB and 18 GB each on the SXM versions, and NVIDIA’s Blackwell datasheet also lists seven for the B200.
- System totals differ slightly. A DGX B200 lists 1,440 GB and 64 TB/s across eight GPUs, so its bandwidth works out a little above the 7.7 TB/s per GPU in the HGX B200 datasheet.
- No list prices for the data-centre parts. NVIDIA’s H100, H200 and DGX B200 pages show no prices, so budgets start from partner quotes. The RTX PRO 6000 is the exception: NVIDIA’s US marketplace listed it at $16,000 (list, verified October 2026).
Bandwidth sets the speed ceiling
Generating a token means reading every active weight from memory once, plus that conversation’s KV cache. One stream therefore cannot decode faster than memory bandwidth ÷ bytes read per token. For a dense model at a short context, the bytes read are close to the size of the weights.
Ceilings for one stream, in tokens per second, from the bandwidth figures above and published parameter counts. Weights use the same arithmetic as our VRAM calculator: 8 bits per weight for FP8 and about 4.25 for AWQ or GPTQ 4-bit.
| GPU | Llama 3.3 70B, FP8 (70.6 GB) | Llama 3.3 70B, 4-bit (37.5 GB) | Qwen3-32B, FP8 (32.8 GB) | Qwen3-32B, 4-bit (17.4 GB) |
|---|---|---|---|---|
| RTX 5090 | Does not fit | Does not fit | Does not fit | ≈ 103 |
| RTX PRO 6000 Workstation | ≈ 25 | ≈ 48 | ≈ 55 | ≈ 103 |
| H100 SXM | ≈ 47, with almost no room for KV cache | ≈ 89 | ≈ 102 | ≈ 192 |
| H200 SXM | ≈ 68 | ≈ 128 | ≈ 147 | ≈ 276 |
| B200 | ≈ 109 | ≈ 205 | ≈ 235 | ≈ 442 |
Three things keep real numbers below these ceilings. Kernels never reach the full rated bandwidth, the KV cache adds bytes to every step as a conversation grows, and serving engines spend time scheduling. The useful signal is the ratio between GPUs: for the same model, a B200 can stream about four times as fast as an RTX PRO 6000, and an H200 about 1.4 times as fast as an H100 SXM.
Batching shifts the picture. A serving engine decodes many sequences in one step, and the weights are read once per step for all of them. With five 8,192-token sessions of Llama 3.3 70B at FP8 on an RTX PRO 6000, a full step reads about 70.6 GB of weights plus up to 13.4 GB of cache. The aggregate ceiling is then about 5 × 1,792 ÷ 84 ≈ 107 tokens per second, shared by the five users. At larger batches the work becomes limited by compute, which is where the FP8 and FP4 rows of the table start to count.
Mixture-of-experts models read only their active experts for each token. gpt-oss-120b keeps 117 billion parameters in memory but activates 5.1 billion per token, so its decode ceiling sits far above a dense model of the same footprint.
How many sessions fit on each GPU
Speed is one half of the sizing question; the other is how many conversations stay resident at once. The method matches our GPU sizing guide:
- KV cache per token = 2 × layers × KV heads × head dimension × bytes per value. Llama 3.3 70B has 80 layers with 8 KV heads of dimension 128, so a 16-bit cache costs 327,680 bytes per token, or 2.68 GB for an 8,192-token session. Qwen3-32B’s 64 layers cost 262,144 bytes per token, or 2.15 GB per session.
- Sessions that fit = (GPU memory × 0.9 − weights) ÷ KV cache per session, which leaves a tenth of memory for activations and the runtime.
- Users ≈ sessions ÷ 0.1 to 0.2, since 10 to 20 percent of named users tend to be active at the busiest moment.
8,192-token sessions that fit, with a 16-bit KV cache / an FP8 KV cache:
| GPU | Llama 3.3 70B, FP8 | Llama 3.3 70B, 4-bit | Qwen3-32B, FP8 | gpt-oss-120b as shipped (65.2 GB) |
|---|---|---|---|---|
| RTX 5090, 32 GB | Does not fit | Does not fit | Does not fit | Does not fit |
| RTX PRO 6000, 96 GB | 5 / 11 | 18 / 36 | 24 / 49 | 68 / 137 |
| H100 SXM, 80 GB | 0 / 1 | 12 / 25 | 18 / 36 | 22 / 44 |
| H100 NVL, 94 GB | 5 / 10 | 17 / 35 | 24 / 48 | 63 / 126 |
| H200, 141 GB | 20 / 41 | 33 / 66 | 43 / 87 | 201 / 402 |
| B200, 180 GB | 34 / 68 | 46 / 92 | 60 / 120 | 315 / 630 |
| B300, 270 GB | 64 / 128 | 76 / 153 | 97 / 195 | 579 / 1,159 |
The gpt-oss-120b column uses the checkpoint as published, with MXFP4 experts. Only half of its 36 layers keep a full-length KV cache; the other half keep a 128-token sliding window. Context length moves every cell: at 32,768 tokens, Llama 3.3 70B at FP8 fits 1 session on an RTX PRO 6000, 5 on an H200 and 8 on a B200. To try other models and contexts, use the calculator.
Two patterns stand out. At FP8, a 70B model leaves an 80 GB H100 SXM with no room for conversations, which is why so many 70B deployments run 4-bit weights or move to H200. And the RTX PRO 6000 sits between H100 and H200 on capacity, so it serves gpt-oss-120b to dozens of sessions even though each stream is slower.
RTX PRO 6000 vs RTX 5090
The two cards share a memory bandwidth figure, 1,792 GB/s, so the question is almost entirely about capacity and operations:
- Memory. 96 GB with error correction against 32 GB without it. On the 5090, a 32B model at 4-bit holds four or five 8,192-token sessions; the RTX PRO 6000 holds about thirty, or a 70B model.
- Compute. The RTX PRO 6000 lists 24,064 CUDA cores and 4,000 sparse FP4 TOPS. The RTX 5090 lists 21,760 CUDA cores and 3,352 sparse FP4 TFLOPS.
- Sharing. MIG on the RTX PRO 6000 splits the card into up to four isolated 24 GB instances, which lets several teams or small models share one GPU. The 5090 has no MIG.
- Power and form. 600 W and 575 W respectively. The 300 W Max-Q version keeps the full 96 GB and bandwidth for multi-GPU workstations, and the Server Edition comes as a dual-slot air-cooled or single-slot liquid-cooled card at 1,597 GB/s.
- Price. NVIDIA’s US list price for the RTX 5090 Founders Edition is $1,999, and its marketplace listed the RTX PRO 6000 at $16,000 (list, verified October 2026). Street prices for the 5090 ran far above list in 2026: Hardware Busters reported Micro Center asking around $4,200 in September.
If one person runs models that fit in 32 GB, the 5090 delivers the same per-stream ceiling for a fraction of the price. Shared serving, 70B-class models and anything that needs ECC or isolation points to the RTX PRO 6000. Our GPU buyer’s guide for local models compares the consumer cards in more detail.
H200 vs H100
Same architecture, different memory. On the SXM versions NVIDIA lists identical Tensor Core peaks, 3,958 sparse FP8 TFLOPS, and the same 700 W ceiling. What changes:
- Capacity: 141 GB against 80 GB, which NVIDIA describes as nearly double.
- Bandwidth: 4.8 TB/s against 3.35 TB/s, about 1.4 times higher, so the per-stream ceiling rises by the same factor.
- MIG slices: seven instances of 18 GB on H200 against seven of 10 GB on H100.
- PCIe versions: H200 NVL keeps 141 GB and 4.8 TB/s at up to 600 W; H100 NVL has 94 GB and 3.9 TB/s at 350–400 W.
NVIDIA’s H200 page claims 1.9 times the Llama 2 70B inference throughput of an H100. The footnote explains much of it: the test ran at batch size 8 on the H100 and batch size 32 on the H200, with 2,000 input and 128 output tokens. More memory means bigger batches, and bigger batches mean more tokens per GPU. That is the case for H200 in one sentence. If your H100s serve 70B-class models with only a handful of sessions each, H200 is the like-for-like upgrade within the same HGX server design.
H200 vs B200, and where B300 fits
The B200 is a larger step. Per GPU, from NVIDIA’s Blackwell datasheet:
- Memory and bandwidth: 180 GB at 7.7 TB/s, against 141 GB at 4.8 TB/s.
- Tensor Cores: 9 PFLOPS of sparse FP8, against 3.958, and FP4 at 18 PFLOPS sparse or 9 dense, which Hopper does not list.
- NVLink: fifth generation at 1.8 TB/s per GPU, twice the H200’s 900 GB/s.
- Power: configurable up to 1,000 W per GPU. A DGX B200 is a 10 RU system rated at about 14.3 kW.
NVIDIA’s DGX B200 page claims 15 times the inference performance of DGX H100, as a projected per-GPU figure for real-time serving with 50 ms between tokens, 32,768 input tokens and 1,028 output tokens. Treat it as a ceiling for that workload, not a general multiplier.
The B300, built on Blackwell Ultra, keeps the 7.7 TB/s and the 9 sparse FP8 PFLOPS but raises memory to 270 GB and dense FP4 to 14 PFLOPS, at up to 1,100 W per GPU. A DGX B300 is rated at about 14 kW and moves to 800 Gb/s ConnectX-8 networking. For memory-bound serving of very large models, B300’s extra 90 GB per GPU matters more than its compute. NVIDIA’s HGX page already lists the next generation, Rubin, at 288 GB of HBM4 and 22 TB/s per GPU.
Decision table: PCIe workstation card or SXM data-centre GPU
| Situation | Better fit | Reason |
|---|---|---|
| One model up to about 80 GB, a few dozen named users | RTX PRO 6000 in a PCIe server | 96 GB per card without SXM boards or liquid cooling |
| Several small models or teams that must stay isolated | RTX PRO 6000 with MIG | Up to four instances with their own memory |
| Racks limited to air cooling and modest power per server | RTX PRO 6000 Max-Q, H100 NVL or H200 NVL | 300 W, 350–400 W or up to 600 W per card |
| A model larger than one GPU, split with tensor parallelism | H200 or B200 on HGX | NVLink at 900 GB/s or 1.8 TB/s; NVIDIA lists no NVLink for the RTX PRO 6000 |
| Hundreds of concurrent or long-context sessions | H200, B200 or B300 | 141 to 270 GB per GPU and 4.8 to 7.7 TB/s |
| FP4 serving at data-centre scale | B200 or B300 | 18 PFLOPS of sparse FP4 per GPU |
| A first pilot on a single server | RTX PRO 6000 or H100 NVL | Standard PCIe hosts, smaller power envelope |
The pattern behind the table: PCIe cards win when a model fits on one card and the facility is built for ordinary servers. SXM GPUs win when the model, the context or the user count outgrows one card, because NVLink keeps multi-GPU serving efficient. Facilities often decide before models do. Eight 1,000 W GPUs plus their host need data-centre power and cooling, which the on-prem server guide covers, and the buy, lease or colocate analysis helps when your own rooms cannot take that load.
How VDF AI fits
VDF AI’s self-hosted package runs as containers, and its documented sizing for the platform itself lists CPU, memory and storage: 4 to 8 vCPUs and 16 to 32 GB of memory for a single-host pilot. The GPUs compared here serve the models; the platform does not consume them.
VDF AI Router keeps applications independent of that hardware choice. Ollama and custom on-premises deployments sit in its model registry next to any cloud models your policy permits, and the router prefers local models when they are available. An RTX PRO 6000 pilot endpoint can later be joined or replaced by H200 or B200 endpoints without changing the applications that call the router. In air-gap mode, routing is restricted to local models only.
Sources
Verified 6 October 2026.
- NVIDIA RTX PRO 6000 Blackwell Workstation Edition
- RTX PRO 6000 Blackwell Workstation Edition datasheet
- NVIDIA RTX PRO 6000 Blackwell Max-Q
- NVIDIA RTX PRO 6000 Blackwell Server Edition
- NVIDIA H100 specifications
- NVIDIA H200 specifications and performance notes
- NVIDIA HGX platform specifications
- NVIDIA Blackwell datasheet
- NVIDIA Blackwell Ultra datasheet
- NVIDIA DGX B200
- NVIDIA DGX B300
- NVIDIA RTX Blackwell GPU architecture whitepaper
- NVIDIA: GeForce RTX 50 Series pricing and availability
- TechSpot: RTX PRO 6000 Blackwell listed at $16,000
- Thunder Compute: RTX PRO 6000 pricing, reviewed 1 October 2026
- Hardware Busters: RTX 50 Founders Edition at MSRP, September 2026
- Llama 3.3 70B config.json (public copy of the gated file)
- Qwen3-32B config.json
- gpt-oss-120b model card
Sizing GPUs for a private model service? Start with the reference architecture or book a demo.