The best GPU for local LLM work is the card whose memory holds your model and context at the highest bandwidth you can afford. In October 2026 that means a 16 GB RTX 5060 Ti or 5070 Ti for 8–14B models, a 32 GB RTX 5090, Radeon AI PRO R9700 or Arc Pro B70 for 27–32B models, and a 96 GB RTX PRO 6000 for 70B-class models.
Two numbers decide whether a GPU suits local language models. Memory capacity decides which models load at all, and memory bandwidth decides how quickly they answer. FP4 support, power draw and software maturity matter once those two are settled.
This guide compares current consumer, workstation and unified-memory options from NVIDIA, AMD, Intel and Apple, using the vendors’ own spec sheets (verified October 2026). List prices come from vendors and their launch announcements; street prices in 2026 often ran far higher. The detailed arithmetic for how much memory a deployment needs lives in our GPU sizing guide, and the VRAM calculator runs it for any model and context.
Quick picks by budget
| Budget at list price | Pick | Memory and bandwidth | Runs well |
|---|---|---|---|
| About $430 | RTX 5060 Ti 16 GB | 16 GB at 448 GB/s | 8–14B models at 4-bit; gpt-oss-20b with short contexts |
| $750–$950 | RTX 5070 Ti for speed, or Arc Pro B70 for capacity | 16 GB at 896 GB/s, or 32 GB at 608 GB/s | 8–14B models faster; 27–32B models on the B70 |
| About $1,300 | Radeon AI PRO R9700 | 32 GB at 640 GB/s | 27–32B models at 4-bit, on ROCm |
| About $2,000 | RTX 5090 | 32 GB at 1,792 GB/s | The fastest single card for models up to about 32B |
| Used market | RTX 3090 or RTX 4090 | 24 GB at 936 or 1,008 GB/s | 24 GB of memory, without FP4 |
| Workstation budget | RTX PRO 6000 Blackwell | 96 GB at 1,792 GB/s | 70B-class models and gpt-oss-120b on one card |
What matters on a GPU spec sheet for LLMs
Memory capacity. Weights, the KV cache for every active conversation and about a tenth spare for the runtime all have to fit. A 14B model at the popular Q4_K_M GGUF quantization, which llama.cpp measures at 4.89 bits per weight, takes 8.5 GB; a 70B model takes 43.1 GB. Context adds more. Ministral 3 14B stores 160 KiB of KV cache per token at 16 bits, so a single 32,768-token conversation adds another 5.4 GB.
Memory bandwidth. Every generated token reads the active weights once, so one stream cannot outrun bandwidth ÷ weight bytes. For that 8.5 GB model the ceiling is about 53 tokens per second on an RTX 5060 Ti, 105 on an RTX 5070 Ti and 210 on an RTX 5090. Real throughput stays below these bounds, but the ratios between cards hold, which is why two 16 GB cards can differ twofold in speed.
FP8 and FP4. Low-precision Tensor Cores speed up prompt processing and let a runtime serve FP8 or NVFP4 checkpoints natively. NVIDIA’s Blackwell generation, covering the RTX 50 series, RTX PRO Blackwell and DGX Spark, supports FP4. The RTX 4090 supports FP8 without FP4, and the RTX 3090 supports neither. AMD lists FP8 and INT4 rates for the RX 9070 XT and Radeon AI PRO R9700 but no FP4, and Intel lists INT8 for the Arc Pro B70.
Power and slots. Board power ranges from 145 W for the single-slot RTX PRO 4000 to 600 W for the RTX PRO 6000 Workstation Edition. NVIDIA specifies a 1,000 W power supply for an RTX 5090 system, so a second high-power card needs a supply and airflow to match.
Software. CUDA has the widest coverage. vLLM’s AMD build lists Instinct MI200 to MI350 plus Radeon RX 7900 and RX 9000-series GPUs, and its Intel build lists Data Center and Arc GPUs. llama.cpp adds HIP for AMD, SYCL for Intel, Vulkan for most GPUs and Metal for Apple silicon. Our comparison of vLLM, Ollama and llama.cpp explains which runtime suits which setup.
Consumer, workstation and unified-memory GPUs compared
Specifications from NVIDIA, AMD, Intel and Apple (verified October 2026). The last column is the bandwidth-bound ceiling for one stream of a 14B model at Q4_K_M (8.5 GB), in tokens per second.
| GPU | VRAM | Bandwidth | FP8 / FP4 | Power | List price | 14B ceiling |
|---|---|---|---|---|---|---|
| RTX 5060 Ti 16 GB | 16 GB GDDR7 | 448 GB/s | Yes / Yes | 180 W | $429 | ≈ 53 |
| RTX 5070 Ti | 16 GB GDDR7 | 896 GB/s | Yes / Yes | 300 W | $749 | ≈ 105 |
| RTX 5080 | 16 GB GDDR7 | 960 GB/s | Yes / Yes | 360 W | $999 | ≈ 113 |
| RTX 5090 | 32 GB GDDR7 | 1,792 GB/s | Yes / Yes | 575 W | $1,999 | ≈ 210 |
| RTX 3090, used | 24 GB GDDR6X | 936 GB/s | No / No | 350 W | None; discontinued | ≈ 110 |
| RTX 4090, used | 24 GB GDDR6X | 1,008 GB/s | Yes / No | 450 W | None; discontinued | ≈ 118 |
| Radeon RX 9070 XT | 16 GB GDDR6 | 640 GB/s | Yes / No | 304 W | $599 | ≈ 75 |
| Radeon AI PRO R9700 | 32 GB GDDR6 | 640 GB/s | Yes / No (INT4 listed) | 300 W | $1,299 | ≈ 75 |
| Intel Arc Pro B70 | 32 GB GDDR6 | 608 GB/s | INT8 listed; FP8 and FP4 not listed | 230 W | $949 | ≈ 71 |
| RTX PRO 4000 Blackwell | 24 GB GDDR7, ECC | 672 GB/s | Yes / Yes | 145 W | Not published by NVIDIA | ≈ 79 |
| RTX PRO 5000 Blackwell | 48 or 72 GB GDDR7, ECC | 1,344 GB/s | Yes / Yes | 300 W | Not published by NVIDIA | ≈ 158 |
| RTX PRO 6000 Blackwell | 96 GB GDDR7, ECC | 1,792 GB/s | Yes / Yes | 600 W; 300 W Max-Q | $16,000 | ≈ 210 |
| DGX Spark | 64 or 128 GB unified | 273 GB/s | FP4 listed | 240 W supply | From $4,999 (64 GB) | ≈ 32 |
| Mac Studio, M5 Max | 36–128 GB unified | 460 or 614 GB/s | Not published | 480 W maximum | From $2,499 | ≈ 54–72 |
| Mac Studio, M5 Ultra | 96–512 GB unified | 1.2 TB/s | Not published | 480 W maximum | From $5,499 | ≈ 141 |
Notes behind the table:
- Where the list prices come from: NVIDIA’s launch announcements for the RTX 5060 Ti, 5070 Ti, 5080 and 5090; AMD’s announced prices for the RX 9070 XT and Radeon AI PRO R9700; Intel’s suggested starting price for the Arc Pro B70; NVIDIA’s US marketplace for the RTX PRO 6000; NVIDIA’s partner launch of the 64 GB DGX Spark; and Apple’s Mac Studio announcement (list, verified October 2026).
- Street prices ran higher. Memory shortages pushed 2026 retail prices well above list. Hardware Busters reported Micro Center asking about $4,200 for an RTX 5090 in September, and TechSpot traced the RTX PRO 6000 from just over $7,600 at pre-order in early 2025 to $16,000.
- DGX Spark pricing moved twice. NVIDIA raised the 128 GB Founders Edition from $3,999 to $4,699 in February 2026, and its product page now lists no price for either memory size. The 64 GB model ships from Acer, ASUS, Dell, Gigabyte, HP and MSI from 23 October.
- FP8 and FP4 on RTX PRO cards follow from their fifth-generation Tensor Cores, which NVIDIA’s datasheets describe as supporting FP4. DGX Spark’s figure is up to 1 PFLOP of sparse FP4.
What fits in 16, 24, 32, 48, 96 and 128 GB
Conversations that fit beside the weights, at 32,768 tokens / 8,192 tokens each. The arithmetic is the calculator’s: 90 percent of a card’s memory is usable, 80 percent of a unified-memory system’s, and the KV cache is held at 16 bits. “No fit” means the weights leave no room for a conversation.
| Model and format | 16 GB | 24 GB | 32 GB | 48 GB | 96 GB | 128 GB unified |
|---|---|---|---|---|---|---|
| Qwen3.5-9B, Q4_K_M (5.9 GB) | 7 / 31 | 14 / 58 | 21 / 85 | 34 / 138 | 74 / 299 | 89 / 359 |
| Ministral 3 14B, Q4_K_M (8.5 GB) | 1 / 4 | 2 / 9 | 3 / 15 | 6 / 25 | 14 / 58 | 17 / 69 |
| gpt-oss-20b as shipped (13.8 GB) | 0 / 3 | 9 / 38 | 18 / 73 | 36 / 143 | 89 / 355 | 109 / 433 |
| Qwen3.6-27B, Q4_K_M (17.0 GB) | No fit | 2 / 8 | 5 / 22 | 12 / 48 | 32 / 129 | 39 / 159 |
| Llama 3.3 70B, Q4_K_M (43.1 GB) | No fit | No fit | No fit | No fit | 4 / 16 | 5 / 22 |
| gpt-oss-120b as shipped (65.2 GB) | No fit | No fit | No fit | No fit | 17 / 68 | 30 / 121 |
Three details explain the odd-looking cells:
- Qwen3.5 and Qwen3.6 hold more conversations than their size suggests. Only a quarter of their layers keep a KV cache, about 32 KiB per token for the 9B model, against 160 KiB for Ministral 3 14B. Each sequence also carries a small fixed state the table leaves out.
- gpt-oss-20b fits 16 GB only with short contexts. OpenAI says it runs within 16 GB of memory, and the arithmetic agrees: three 8,192-token conversations, but not one of 32,768.
- Llama 3.3 70B needs a leaner format for 48 GB. At about 4.25 bits per weight (AWQ or GPTQ) it takes 37.5 GB and leaves room for two 8,192-token conversations. A 72 GB RTX PRO 5000 holds it at Q4_K_M with eight. The new 64 GB DGX Spark cannot hold gpt-oss-120b at all.
Best budget GPU for LLM inference
Four routes compete below about $1,300 at list price, and each gives up something different.
- RTX 5060 Ti 16 GB, $429. The lowest list price for a current NVIDIA card with 16 GB, with Blackwell’s FP4 support and a 180 W board power. Its 128-bit bus delivers 448 GB/s, half the RTX 5070 Ti’s 896 GB/s, so the same model streams at half the ceiling. It suits one person running 8–14B models.
- Radeon RX 9070 XT, $599. The same 16 GB, 640 GB/s of bandwidth and FP8, on vLLM’s list of supported RX 9000-series cards. It beats the 5060 Ti on speed, not on capacity.
- 32 GB for under $1,300: Arc Pro B70 or Radeon AI PRO R9700. Both double the memory of any 16 GB card, enough for 27–32B models at 4-bit with room for context. Their 608 and 640 GB/s make each stream slower than on an RTX 5070 Ti. The R9700 adds FP8 and error-correcting memory on Linux; the B70 draws 230 W against the R9700’s 300 W.
- A used RTX 3090 or 4090. 24 GB at 936 or 1,008 GB/s, without a vendor list price or warranty. The 3090 has no FP8 or FP4 Tensor Cores, and the pair draws 350 and 450 W.
Buy for memory first. A slower card that holds the model beats a faster one that cannot load it, and the gap between budget cards is smaller than the gap between a model that fits and one that does not.
Best consumer GPU for LLM inference
The RTX 5090 is the fastest consumer card for local models: 32 GB of GDDR7 on a 512-bit bus at 1,792 GB/s, with 3,352 sparse FP4 TFLOPS from its fifth-generation Tensor Cores. It has nearly twice the bandwidth of an RTX 5080, which is why its ceiling for the 14B example reaches about 210 tokens per second. Its 32 GB runs a 27B model at 4-bit with several long conversations, or a 9B model with dozens.
Between the 16 GB cards, the RTX 5070 Ti is the better buy for language models. It has 93 percent of the RTX 5080’s bandwidth, 896 GB/s against 960, at 75 percent of its $999 list price. The extra compute in the 5080 does little for single-user chat, where memory reads set the pace.
A used RTX 4090 still out-streams every 16 GB Blackwell card, with 1,008 GB/s, and adds 8 GB of capacity. It lacks FP4, which matters only if you plan to serve NVFP4 checkpoints such as those NVIDIA and Mistral now publish on Hugging Face.
Workstation, data-centre and unified-memory options
RTX PRO Blackwell. The 24 GB RTX PRO 4000 is a single-slot, 145 W card with ECC, built for compact workstations. The RTX PRO 5000 comes with 48 or 72 GB at 1,344 GB/s and 300 W, and splits into two MIG instances. The RTX PRO 6000 carries 96 GB at 1,792 GB/s, splits into up to four MIG instances, and comes as a 600 W Workstation Edition, a 300 W Max-Q card with the same memory and bandwidth, and a Server Edition. These cards bring ECC memory, MIG isolation and capacity that consumer cards lack. Our RTX PRO 6000, H100, H200 and B200 comparison covers them against data-centre GPUs.
Data-centre GPUs. For shared serving, NVIDIA’s H100 (80 GB at 3.35 TB/s), H200 (141 GB at 4.8 TB/s) and B200 (180 GB at 7.7 TB/s) and AMD’s Instinct MI350X and MI355X (288 GB of HBM3E at 8 TB/s, at 1,000 and 1,400 W) carry no list prices on the vendors’ spec pages, so budgets start from partner quotes. Choosing the server around them is the subject of our server sizing guide.
Unified memory. DGX Spark pairs 64 or 128 GB with 273 GB/s and NVIDIA’s CUDA stack, and NVIDIA says two 64 GB units cabled together pool 128 GB. Apple’s Mac Studio offers the M5 Max with up to 128 GB at 460 or 614 GB/s from $2,499, and the M5 Ultra with 1.2 TB/s and up to 256 GB from $5,499, with a 512 GB option due in late October. These systems load models no single consumer card can, at lower speed per stream. The DGX Spark, Mac Studio and Ryzen AI Max comparison goes deeper.
Recommendations by model size
| Models you plan to run | Plan for (4-bit weights, one long conversation) | Budget pick | Comfortable pick |
|---|---|---|---|
| 3–9B, such as Qwen3.5-9B, Granite 4.2 8B or Ministral 3 8B | 8–12 GB | RTX 5060 Ti 16 GB | RTX 5070 Ti |
| 12–20B, such as Gemma 4 12B, Ministral 3 14B or gpt-oss-20b | 16–24 GB | RTX 5060 Ti 16 GB, short contexts | RTX 5090 or a used RTX 4090 |
| 27–32B, such as Qwen3.6-27B, Gemma 4 31B or Qwen3-32B | 24–32 GB | Arc Pro B70 or a used RTX 3090 | RTX 5090 or Radeon AI PRO R9700 |
| 70B dense, such as Llama 3.3 70B | 48–72 GB | RTX PRO 5000 48 GB at 4.25 bits per weight | RTX PRO 5000 72 GB or RTX PRO 6000 |
| 100B-plus mixture of experts, such as gpt-oss-120b | 96–128 GB | DGX Spark 128 GB or Mac Studio, at lower speed | RTX PRO 6000 |
Most teams do not need the largest card. Small models now handle tool calling, retrieval and image input well, and our small language model shortlist matches them to tasks. Start from the model your workload needs, then buy the cheapest card that holds it with room for your longest conversations. The VRAM tier guide lists which models fit each card, and the Mac model guide does the same for unified memory.
How VDF AI fits
VDF AI’s own services do not need a GPU: the self-hosting documentation sizes a pilot host by CPU, memory and storage alone. The cards in this guide serve the models, through Ollama or another on-premises deployment registered with the platform.
VDF AI Router sends small tasks to small models, so a 16 GB card running a 9B model can absorb routine requests while a larger GPU takes the hard ones. It probes local runtimes continuously. Outside air-gap mode, it shifts traffic to permitted cloud models if a local endpoint goes down; in air-gap mode it routes to local models only.
Sources
Verified 6 October 2026.
- NVIDIA RTX Blackwell GPU architecture whitepaper (RTX 5090, 5080, 5070 Ti, 4090 and 3090 specifications)
- NVIDIA GeForce RTX 5090 specifications
- NVIDIA GeForce RTX 5060 family specifications
- NVIDIA: RTX 5060 family pricing and availability
- PC Perspective: RTX 5060 Ti 16 GB memory specifications
- NVIDIA: GeForce RTX 50 Series pricing and availability
- Hardware Busters: RTX 50 Founders Edition at MSRP, September 2026
- AMD Radeon AI PRO R9700 specifications
- Phoronix: Radeon AI PRO R9700 official price
- AMD Radeon RX 9070 XT specifications
- TechRadar: RX 9070 XT launch price
- Intel Arc Pro B70 specifications
- Intel Arc Pro B-Series graphics cards
- Phoronix: Intel Arc Pro B70 announcement and price
- NVIDIA RTX PRO 4000 Blackwell
- NVIDIA RTX PRO 5000 Blackwell
- NVIDIA RTX PRO 6000 Blackwell Workstation Edition
- NVIDIA RTX PRO 6000 Blackwell Max-Q
- TechSpot: RTX PRO 6000 Blackwell listed at $16,000
- Thunder Compute: RTX PRO 6000 pricing, reviewed 1 October 2026
- NVIDIA DGX Spark specifications
- NVIDIA: DGX Spark 64 GB announcement
- NVIDIA: DGX Spark price change, February 2026
- Apple Mac Studio technical specifications
- Apple: Mac Studio with M5 Max and M5 Ultra
- AMD Instinct MI350X and MI355X
- NVIDIA H100, H200 and HGX B200
- vLLM GPU installation: AMD ROCm and Intel XPU
- llama.cpp supported backends
- llama.cpp quantize README: Q4_K_M bits per weight
- Ministral 3 14B config.json, Qwen3.5-9B model card and Qwen3.6-27B config.json
- gpt-oss-20b and gpt-oss-120b model cards
- NVIDIA Qwen3.8-27B NVFP4 checkpoint and Mistral Small 4 NVFP4 checkpoint
Planning GPUs for a team rather than one desk? See the reference architecture or book a demo.