The best local LLM for a Mac is decided by unified memory. On a 16 GB MacBook Air or Mac mini, run Qwen3.5-9B or Gemma 4 12B at 4-bit; at 32 to 36 GB, Qwen3.6-35B-A3B or Qwen3.8-27B; at 64 GB, Llama 3.3 70B fits; at 128 GB, gpt-oss-120b or Qwen3.5-122B-A10B. Only a 512 GB Mac Studio holds Qwen3.5-397B-A17B.
Apple silicon Macs have one pool of unified memory that the CPU and GPU share. That makes the memory you ordered the hard limit on model size, and memory bandwidth the main limit on speed. Apple’s October 2026 line-up runs from 16 GB to 512 GB, so the right model differs by more than an order of magnitude from one Mac to the next.
Memory options and bandwidth below come from Apple’s technical specification pages, model facts from each vendor’s Hugging Face card, and file sizes from the MLX conversions published on Hugging Face, all checked on 6 October 2026. For a hardware comparison against NVIDIA and AMD desk-side systems, see our DGX Spark, Mac Studio and Ryzen AI Max comparison.
Quick picks by Mac
| Mac (October 2026) | Memory you can order | Best local LLM picks |
|---|---|---|
| MacBook Air (M5) | 16, 24 or 32 GB | Qwen3.5-9B or Gemma 4 12B; Gemma 4 26B A4B from 24 GB |
| Mac mini (M6), from $899 | 16, 24 or 32 GB | Same as the Air; Qwen3.6-35B-A3B at 32 GB |
| Mac mini (M5 Pro), from $1,699 | 24, 48 or 64 GB | Qwen3.8-27B at 8-bit (48 GB); Llama 3.3 70B at 4-bit (64 GB) |
| MacBook Pro (M5) | 16, 24 or 32 GB | Qwen3.5-9B, Gemma 4 12B, Gemma 4 26B A4B |
| MacBook Pro (M5 Pro or M5 Max) | 24, 36, 48, 64 or 128 GB | Qwen3.6-35B-A3B; gpt-oss-120b at 128 GB |
| Mac Studio (M5 Max), from $2,499 | 36, 48, 64 or 128 GB | Gemma 4 31B, Qwen3.6-35B-A3B; Qwen3.5-122B-A10B at 128 GB |
| Mac Studio (M5 Ultra), from $5,499 | 96, 256 or 512 GB | Qwen3.5-122B-A10B; Qwen3.5-397B-A17B at 512 GB |
| iMac (M4) | 16, 24 or 32 GB | The 16–32 GB picks above |
Prices are Apple’s US starting prices from its August and September 2026 announcements (list, verified October 2026); memory upgrades cost extra.
Mac memory tiers in October 2026
| Mac | Chip | Unified memory options | Memory bandwidth |
|---|---|---|---|
| MacBook Air 13 and 15-inch | M5 | 16, 24, 32 GB | 153 GB/s |
| Mac mini | M6 | 16 GB; 24 or 32 GB | 153 GB/s; 170 GB/s with 24 or 32 GB |
| Mac mini | M5 Pro | 24, 48, 64 GB | 307 GB/s |
| MacBook Pro 14-inch | M5 | 16, 24, 32 GB | 153 GB/s |
| MacBook Pro 14 and 16-inch | M5 Pro | 24, 48, 64 GB | 307 GB/s |
| MacBook Pro 14 and 16-inch | M5 Max, 32-core GPU | 36 GB | 460 GB/s |
| MacBook Pro 14 and 16-inch | M5 Max, 40-core GPU | 48, 64, 128 GB | 614 GB/s |
| Mac Studio | M5 Max | 36 GB (32-core GPU); 48, 64, 128 GB (40-core GPU) | 460 or 614 GB/s |
| Mac Studio | M5 Ultra | 96 GB; 256 or 512 GB with the 80-core GPU | 1.2 TB/s |
| iMac | M4 | 16, 24, 32 GB | 120 GB/s |
Apple’s September announcement said the 512 GB Mac Studio would follow in late October. The tiers people search for, 16, 24, 32, 48, 64, 128, 256 and 512 GB, all exist, along with 36 GB and 96 GB as base configurations of the M5 Max and M5 Ultra.
How much memory a model can use on a Mac
Our sizing follows the VRAM calculator, which treats 80% of a unified-memory pool as usable because macOS and your apps need the rest. Inside that budget go the weights, a KV cache that grows with every token, and, for hybrid models such as Qwen3.5, a small fixed state per session. The full method is in our GPU sizing guide.
Worked example: Qwen3.6-35B-A3B on a 32 GB Mac.
- Budget: 32 GB × 0.8 = 25.6 GB.
- Weights: the MLX 4-bit conversion is 20.40 GB.
- Fixed state: 30 linear-attention layers hold about 0.06 GB per session.
- KV cache: only 10 of 40 layers keep a growing cache, at 2 × 10 × 2 KV heads × 256 × 2 bytes = 20,480 bytes per token.
- Context: (25.6 − 20.40 − 0.06) GB ÷ 20,480 bytes ≈ 250,000 tokens, almost the model’s 262,144 limit.
The MLX LM documentation adds a practical note. Models that are large relative to RAM can be slow, and on macOS 15 or later you can raise the GPU wired memory limit with sudo sysctl iogpu.wired_limit_mb=N, choosing a value above the model size and below total memory.
Bandwidth decides speed once a model fits. Every generated token reads the active weights, so a mixture-of-experts model with 3B active parameters reads roughly a ninth of the bytes that a 27B dense model reads. That is why MoE models suit the 153 to 170 GB/s chips, and why dense 27B to 70B models feel better on M5 Max and M5 Ultra.
What fits in each memory tier
All sizes are the MLX 4-bit or native-format conversions on Hugging Face. Context is for one session with a 16-bit cache, capped at each model’s limit.
| Memory | Usable (80%) | Models that fit, with weights | Context for one chat |
|---|---|---|---|
| 16 GB | 12.8 GB | Qwen3.5-9B 6.0 GB; Gemma 4 12B 6.7 GB; gpt-oss-20b 12.1 GB | ≈207K; 256K; ≈29K |
| 24 GB | 19.2 GB | Gemma 4 26B A4B 15.3 GB; gpt-oss-20b; Qwen3.8-27B 16.1 GB | ≈178K; 128K; ≈46K |
| 32 GB | 25.6 GB | Qwen3.6-35B-A3B 20.4 GB; Qwen3.8-27B; Gemma 4 31B 18.4 GB | ≈250K; ≈143K; ≈78K |
| 36 GB | 28.8 GB | Qwen3.6-35B-A3B; Qwen3.8-27B; Gemma 4 31B | 256K; ≈192K; ≈117K |
| 48 GB | 38.4 GB | Qwen3.8-27B at 8-bit 29.5 GB; Gemma 4 31B at 4-bit | ≈133K; ≈234K |
| 64 GB | 51.2 GB | Llama 3.3 70B 39.7 GB; Qwen3.6-35B-A3B at 8-bit 37.7 GB; Gemma 4 31B at 8-bit 33.8 GB | ≈35K; 256K; ≈203K |
| 96 GB | 76.8 GB | gpt-oss-120b 63.4 GB; Qwen3.5-122B-A10B 69.6 GB; Mistral Small 4 67.8 GB | 128K; 256K; 256K |
| 128 GB | 102.4 GB | The 96 GB models with room for four long sessions; Llama 3.3 70B at 8-bit 75.0 GB | ≈84K for the 8-bit 70B |
| 256 GB | 204.8 GB | Qwen3.5-122B-A10B at 8-bit 130.7 GB; DeepSeek-V4-Flash at 4-bit 151.5 GB | 256K; see the card |
| 512 GB | 409.6 GB | Qwen3.5-397B-A17B at 4-bit 223.9 GB; GLM-5.3-Flash at 4-bit 204.0 GB | 256K for Qwen; see the card for GLM |
Two near misses are worth knowing. Gemma 4 31B at 4-bit loads on a 24 GB Mac but leaves under 1,000 tokens of context, and GLM-5.3-Flash at 4-bit sits right at a 256 GB Mac’s budget with nothing left for a cache. DeepSeek-V4 and GLM-5.3 use compressed or sparse attention that our per-token formula does not cover, so read their cards before sizing long contexts.
Best local LLM for Mac mini
Apple sells two Mac minis. The M6 model starts at 16 GB with 153 GB/s and goes to 24 or 32 GB at 170 GB/s. The M5 Pro model offers 24, 48 or 64 GB with 307 GB/s.
- Mac mini M6, 16 GB. Qwen3.5-9B is the best all-rounder; Qwen reports 66.1 on BFCL-V4 and 63.0 on AA-LCR for it. Gemma 4 12B adds audio and image input and keeps its full 256K window. gpt-oss-20b loads, but with about 29,000 tokens of context.
- Mac mini M6, 24 or 32 GB. Gemma 4 26B A4B at 24 GB; Qwen3.6-35B-A3B at 32 GB. Both activate only 3 to 4 billion parameters per token, which matters on a 170 GB/s chip.
- Mac mini M5 Pro, 48 or 64 GB. The M5 Pro’s doubled bandwidth makes dense models usable. At 48 GB run Qwen3.8-27B at 8-bit; at 64 GB, Llama 3.3 70B at 4-bit, under Meta’s Llama 3.3 Community License.
A Mac mini also makes a quiet always-on endpoint for one person or a small team, with the caveats in the runtime section below.
Best local LLM for MacBook Pro
The MacBook Pro spans more memory than any other Mac line: 16 GB with the M5 up to 128 GB with the 40-core M5 Max.
- M5, 16 to 32 GB. The same picks as the MacBook Air: Qwen3.5-9B and Gemma 4 12B, with Gemma 4 26B A4B from 24 GB.
- M5 Pro, 24 to 64 GB. At 48 GB, Gemma 4 31B at 4-bit keeps about 234,000 tokens of context; at 64 GB, Qwen3.6-35B-A3B at 8-bit keeps its full window.
- M5 Max, 36 GB. The 32-core GPU version comes only with 36 GB at 460 GB/s. Qwen3.6-35B-A3B and Qwen3.8-27B both fit with long context.
- M5 Max, 128 GB. The only laptop that holds gpt-oss-120b (OpenAI’s 117B-parameter MoE with 5.1B active) or Qwen3.5-122B-A10B. Mistral Small 4, a 119B MoE with 6.5B active and a 256K window under Apache 2.0, also fits.
For writing and reviewing code on a laptop, our local coding model guide lists coder-specific checkpoints.
Best LLM for Mac Studio
The Mac Studio is where a Mac stops being a personal machine and starts holding server-class models.
- M5 Max, 36 to 128 GB. At 128 GB and 614 GB/s, gpt-oss-120b, Qwen3.5-122B-A10B and Mistral Small 4 all fit with several long sessions at once.
- M5 Ultra, 96 GB. The base Ultra already fits the 120B class: gpt-oss-120b with its full 128K window, Qwen3.5-122B-A10B with 256K.
- M5 Ultra, 256 GB. Qwen3.5-122B-A10B at 8-bit for higher fidelity, or DeepSeek-V4-Flash, MIT-licensed, at 4-bit.
- M5 Ultra, 512 GB. Needed for Qwen3.5-397B-A17B at 4-bit, the largest Apache 2.0 model in this list, or GLM-5.3-Flash (MIT) with room for context.
Check the licence before you download a model this size. MiniMax M2.7 fits a 256 GB Mac at 4-bit, but its licence permits non-commercial use only and requires written authorization for commercial use. Our local LLM overview covers the other licence families you will meet.
MLX vs llama.cpp, Ollama and LM Studio
| Runtime | What it runs on a Mac | Strengths | Watch for |
|---|---|---|---|
| MLX LM (Apple) | MLX-format safetensors, such as the mlx-community conversions | Apple’s framework; converts and quantizes Hub models; prompt caching; distributed inference across Macs | Its HTTP server is documented as not recommended for production because it only implements basic security checks |
| llama.cpp | GGUF files on the Metal backend | Apple silicon is a first-class target; the same file runs on PCs and servers; JSON-schema output; quantized KV cache | GGUF sizes vary by builder, so check the file |
| Ollama | GGUF through llama.cpp, plus an MLX engine for supported architectures | One install and a model library; a September 2026 pre-release makes MLX the default for supported models such as Qwen3.5 to 3.8 and Gemma 4 | Default context follows GPU memory: 4K below 24 GiB, 32K up to 48 GiB |
| LM Studio | GGUF and MLX; the MLX engine ships with LM Studio 0.3.4 and later | Desktop app with a model browser and local server | Needs Apple silicon and macOS 14 or newer |
Pick by what you will do next. MLX gives the most direct path on the Mac itself. llama.cpp and Ollama are the portable choice if the same model will later run on Linux GPU servers, where our vLLM, Ollama and llama.cpp comparison picks up. For PC graphics cards, see the VRAM tier guide; for retrieval and agent workloads on any hardware, see our picks for local RAG and agents and tool calling. Our Gemma 4 setup guide includes the MLX and Ollama install steps for that family.
How VDF AI fits
VDF AI’s own self-hosted platform is not a Mac workload: its infrastructure requirements list macOS through Docker Desktop for local evaluation and developer environments only, not production. Linux on Apple silicon and other ARM64 hardware is fully supported through multi-architecture images.
What a Mac can do in a VDF AI pilot is host a model. VDF AI Router registers Ollama and custom on-premises deployments alongside approved cloud models, and its air-gap mode keeps routing on local models. When the pilot moves to rack servers, applications keep calling the router while the model endpoint behind it changes.
Sources
Specifications, model cards and file listings verified 6 October 2026.
- Mac mini technical specifications
- MacBook Pro technical specifications
- MacBook Air technical specifications
- Mac Studio technical specifications
- iMac technical specifications
- Apple: Mac mini with M6 and M5 Pro
- Apple: the new Mac mini and Mac Studio are available
- MLX LM and MLX LM server notes
- llama.cpp
- Ollama: MLX on Apple silicon preview and Ollama v0.40.0-rc5 release notes
- Ollama: context length
- LM Studio system requirements and LM Studio mlx-engine
- MLX conversions: Qwen3.6-35B-A3B 4-bit, Qwen3.8-27B 4-bit, Gemma 4 31B 4-bit, Qwen3.5-122B-A10B 4-bit, gpt-oss-120b, Qwen3.5-397B-A17B 4-bit, Llama 3.3 70B 4-bit
- Model cards: Qwen3.5-9B, Qwen3.6-35B-A3B, Gemma 4 31B, gpt-oss-120b, Mistral Small 4, DeepSeek-V4-Flash-0731, GLM-5.3-Flash
- MiniMax M2.7 licence
Piloting local models on a Mac before buying servers? See how VDF AI Router keeps applications on one endpoint, or book a demo.