The best small language models in 2026 are Qwen3.5-9B and Granite 4.2 8B for tool calling and agents, Gemma 4 E4B and Ministral 3 8B for image input on modest hardware, and Phi-4-mini or Llama 3.2 3B where only a 3B footprint fits. All except Llama 3.2 are Apache 2.0 or MIT, and each fits an 8–16 GB GPU at 4-bit.
Small language models changed more in the last year than their parameter counts suggest. Current releases between 1 and 14 billion parameters handle tool calls, long documents and images, and several publish agent benchmarks on their model cards. For a platform team, that changes the default question from “which large model” to “which tasks still need one”.
This guide compares eleven open-weight families. Licences, parameter counts, context lengths and tool-calling notes were checked against the vendors’ model cards on Hugging Face in October 2026. Scores appear only where a vendor or a public leaderboard published them, with the benchmark named. For the definition and the basic trade-offs, see our small language model explainer.
Quick picks by task
| Task | Pick | Why, from the model card |
|---|---|---|
| Tool calling and agents | Qwen3.5-9B | Qwen reports 66.1 on BFCL-V4 and 79.1 on TAU2-Bench; vLLM tool parser documented |
| Tool calling with a dense, text-only model | Granite 4.2 8B | IBM lists tool calling and agentic workflows as primary uses; 128K context, extendable to 512K |
| Retrieval over long documents | Qwen3.5-4B or Qwen3.5-9B | 262,144-token native context with a KV cache in only 8 of 32 layers |
| Text, images and audio on a laptop | Gemma 4 E4B | Google built the E models for laptops and mobile devices |
| Images plus function calling | Ministral 3 8B | 0.4B vision encoder, native function calling and JSON output; fits 12 GB in FP8, per Mistral |
| The strongest model under 15B | Gemma 4 12B or Ministral 3 14B | Google reports 77.2% on MMLU Pro for the 12B; Mistral compares the 14B with its 24B Mistral Small 3.2 |
| Phones and CPUs | Qwen3.5-0.8B, Llama 3.2 1B or LFM2.5-1.2B | Under 1 GB of weights at 4-bit |
| Fully open training data | SmolLM3-3B or Olmo 3 7B | Weights, data mixtures and training details published |
What counts as a small language model in 2026
The working range today is roughly 0.5 to 14 billion parameters, small enough to run on one consumer GPU, a laptop or, at the bottom end, a phone. Two architectures stretch the definition. Mixture-of-experts models such as Liquid AI’s LFM2.5-8B-A1B keep 8.3 billion parameters in memory but activate 1.5 billion per token. Google’s Gemma 4 E2B and E4B count “effective” parameters: 2.3B and 4.5B, or 5.1B and 8B including their per-layer embeddings.
Three things separate the 2026 generation from the 7B models of 2024:
- Built-in vision. Every Qwen3.5 size includes a vision encoder, Ministral 3 adds a 0.4B one, and Gemma 4’s small models accept audio as well as images.
- Long context at low cost. Qwen3.5 and LFM2.5 mix full attention with linear or convolution layers, so only some layers keep a growing KV cache. A 262K-token window on a 9B model is now practical on one GPU.
- Training for tool use. Model cards now report function-calling and agent benchmarks such as BFCL and TAU2, and name the vLLM parser that reads the model’s tool calls.
The shortlist: eleven small model families
Checked against the Hugging Face model cards in October 2026. Parameter counts are the Hub totals, which include vision or audio encoders where a model has them.
| Model (vendor) | Year | Parameters | Context | Inputs | Licence | Tool calling, per the card |
|---|---|---|---|---|---|---|
| Qwen3.5-0.8B, 2B, 4B, 9B (Qwen) | 2026 | 0.9B, 2.3B, 4.7B, 9.7B | 262K | Text, image, video | Apache 2.0 | Yes; vLLM parser qwen3_coder |
| Gemma 4 E2B, E4B (Google) | 2026 | 2.3B and 4.5B effective | 128K | Text, image, audio | Apache 2.0 | Native function calling |
| Gemma 4 12B (Google) | 2026 | 12.0B | 256K | Text, image, audio | Apache 2.0 | Native function calling |
| Ministral 3 3B, 8B, 14B (Mistral) | 2025 | 3.8B, 8.9B, 13.9B | 256K | Text, image | Apache 2.0 | Native function calling and JSON; parser mistral |
| Granite 4.2 3B, 8B (IBM) | 2026 | 3.7B, 8.8B | 128K, extendable to 512K | Text | Apache 2.0 | Reasoning before tool calls; parser qwen3_coder |
| Phi-4-mini-instruct (Microsoft) | 2025 | 3.8B | 128K | Text | MIT | Dedicated function-calling format |
| Llama 3.2 1B, 3B (Meta) | 2024 | 1.2B, 3.2B | 128K | Text | Llama 3.2 Community License | Aimed at agentic retrieval and summarization |
| SmolLM3-3B (Hugging Face) | 2025 | 3.1B | 64K, 128K with YaRN | Text | Apache 2.0 | XML or Python-style calls; parser hermes |
| LFM2.5-1.2B, 2.6B, 8B-A1B (Liquid AI) | 2026 | 1.2B, 2.7B, 8.3B with 1.5B active | 32K for 1.2B; 128K for the others | Text | LFM Open License 1.0 | Pythonic function calls |
| Nemotron 3 Nano 4B (NVIDIA) | 2026 | 4.0B | Up to 262K | Text, English only | NVIDIA Nemotron Open Model License | Yes; parser qwen3_coder |
| MiniCPM5-2B (OpenBMB) | 2026 | 2.5B | 128K | Text | Apache 2.0 | Designed for tool-use workflows |
| Olmo 3 7B Instruct (Ai2) | 2025 | 7.3B | 64K | Text, English | Apache 2.0 | Function-calling system prompt |
A few notes behind the table. Granite 4.2’s card gives 25 August 2026 as the release date, and NVIDIA’s gives 16 March 2026 for Nemotron 3 Nano 4B, which NVIDIA compressed from its 9B Nemotron Nano v2. Qwen lists the 0.8B model’s intended uses as prototyping, task-specific fine-tuning and research. Just above the range, OpenAI’s gpt-oss-20b has 21B parameters with 3.6B active and runs within 16 GB of memory.
Best small language model for tool calling
Vendors now publish function-calling scores, but each runs its own harness and settings. Compare sizes within one family, not numbers across families.
| Family | Benchmark | Smallest → largest size in the family |
|---|---|---|
| Qwen3.5 | BFCL-V4 | 0.8B: 25.3 · 2B: 43.6 · 4B: 50.3 · 9B: 66.1 |
| Qwen3.5 | TAU2-Bench | 0.8B: 11.6 · 2B: 48.8 · 4B: 79.9 · 9B: 79.1 |
| Granite 4.2 | BFCL v4 | 3B: 52.41 · 8B: 52.39 |
| Granite 4.2 | τ³-bench | 3B: 50.99 · 8B: 66.34 |
| Gemma 4 | Tau2, average of three | E2B: 24.5% · E4B: 42.2% · 12B: 69.0% |
Two patterns stand out. In every family that publishes these numbers, scores climb steeply from the sub-1B and 2B sizes to 4B. And Qwen3.5-4B matches the 9B on TAU2-Bench while trailing it clearly on BFCL-V4, so test both when memory is tight.
The independent Berkeley Function Calling Leaderboard, last updated on 12 April 2026, does not yet list the 2026 small models, but its older entries show the same direction. Qwen3-8B in function-calling mode scored 42.57% overall, ahead of Mistral Large 2411 at 38.37%. Llama 3.2 3B scored 21.95% and Llama 3.2 1B 10.82%. Salesforce’s xLAM-2-8b-fc-r reached 46.68%, but its CC-BY-NC-4.0 licence rules out commercial use.
Three settings from the model cards make small models more reliable with tools:
- Use the parser the card names. vLLM needs
--enable-auto-tool-choiceplus the right parser:qwen3_coderfor Qwen3.5, Granite 4.2 and Nemotron 3 Nano 4B,mistralfor Ministral 3 andhermesfor SmolLM3. - Keep the tool list short. Mistral advises limiting tools to the minimum the use case needs, and a temperature below 0.1 in production.
- Follow the sampling advice. NVIDIA recommends a temperature of 0.6 for tool calling with Nemotron 3 Nano 4B, lower than for reasoning tasks.
When a small model is not enough, our agent and tool-calling picks cover the larger models, and the VRAM tier guide shows what each card can hold.
RAG and long documents: context costs memory
Retrieval pipelines put long passages into every prompt, so the KV cache, not the weights, often decides whether a small model fits. Per token at 16 bits it equals 2 × layers with a cache × KV heads × head dimension × 2 bytes, from each model’s config.json:
| Model | Layers with a growing KV cache | KV cache per token | One 32K-token conversation | One 128K-token conversation |
|---|---|---|---|---|
| Qwen3.5-0.8B | 6 of 24 | 12 KiB | 0.4 GB | 1.6 GB |
| Qwen3.5-9B | 8 of 32 | 32 KiB | 1.1 GB | 4.3 GB |
| Llama 3.2 1B | 16 of 16 | 32 KiB | 1.1 GB | 4.3 GB |
| SmolLM3-3B | 36 of 36 | 72 KiB | 2.4 GB | 9.7 GB |
| Llama 3.2 3B | 28 of 28 | 112 KiB | 3.8 GB | 15.0 GB |
| Phi-4-mini | 32 of 32 | 128 KiB | 4.3 GB | 17.2 GB |
| Ministral 3 8B | 34 of 34 | 136 KiB | 4.6 GB | 18.3 GB |
| Granite 4.2 8B | 40 of 40 | 160 KiB | 5.4 GB | 21.5 GB |
The difference is architectural. Granite 4.2 8B is a dense transformer that caches every layer, so at 128K tokens its cache is four times the size of its 4-bit weights. Qwen3.5-9B caches only its 8 full-attention layers; the other 24 carry a fixed-size state instead. For one GPU serving several long-context retrieval sessions, the hybrid design is the easier fit. Liquid AI likewise recommends LFM2.5-2.6B for retrieval and data extraction, but not for knowledge-heavy tasks.
Two caveats. A long context window does not mean a model uses all of it well, so test retrieval quality at the lengths you plan to serve. And retrieval usually beats stuffing whole documents into the prompt; our note on long context versus retrieval explains when each wins.
Hardware and edge deployment
Weights at three precisions, using the same arithmetic as our VRAM calculator: 16 bits, 8 bits and 4.89 bits per weight for GGUF Q4_K_M.
| Size class | Examples | BF16 | FP8 | Q4_K_M | Typical hardware |
|---|---|---|---|---|---|
| Under 1.5B | Qwen3.5-0.8B, Llama 3.2 1B, LFM2.5-1.2B | 1.7–2.5 GB | 0.9–1.2 GB | 0.5–0.8 GB | Phone, CPU or any GPU |
| 2–4B | Qwen3.5-4B, Phi-4-mini, Ministral 3 3B, Granite 4.2 3B, SmolLM3-3B | 6.2–9.3 GB | 3.1–4.7 GB | 1.9–2.8 GB | Laptop or an 8 GB GPU |
| 7–9B | Qwen3.5-9B, Granite 4.2 8B, Ministral 3 8B, Olmo 3 7B | 14.6–19.3 GB | 7.3–9.7 GB | 4.5–5.9 GB | A 12–16 GB GPU |
| 12–14B | Gemma 4 12B, Ministral 3 14B, Phi-4 | 23.9–29.3 GB | 12.0–14.7 GB | 7.3–9.0 GB | A 16–24 GB GPU |
Add the KV cache from the previous table for each conversation you serve at once, plus about a tenth of memory for the runtime. A 16 GB card holds a 9B model at 4-bit with several long conversations; our GPU buyer’s guide compares the cards by budget.
Small models with vision, audio and edge targets
Several small models now read images, which removes a separate vision service from many pipelines.
- Qwen3.5 accepts image and video input at every size from 0.8B up. Qwen’s tables show the 9B ahead of its own previous Qwen3-VL-30B-A3B on most vision benchmarks it reports.
- Gemma 4 E2B and E4B take text, images and audio, use 512-token sliding-window layers to save memory, and let you trade detail for speed with visual token budgets from 70 to 1,120 tokens per image.
- Ministral 3 pairs each model with a 0.4B vision encoder. Mistral says the 3B fits in 8 GB of VRAM at FP8, the 8B in 12 GB and the 14B in 24 GB.
- MiniCPM-V 4.6 builds a 1.3B vision model on Qwen3.5-0.8B, and OpenBMB demonstrates it on iOS, Android and HarmonyOS phones.
- Nemotron 3 Nano 4B targets edge platforms such as Jetson Thor, GeForce RTX cards and DGX Spark, in English only.
Our guide to the best local vision models compares the larger vision models as well.
Small language models vs large language models
Small models win on narrow, repeated work: classification, extraction, routing, short answers grounded in retrieved text and well-defined tool calls. They answer faster, fit on hardware you already own and can be fine-tuned on your own examples. The published evidence for how close they come:
- Qwen’s own table puts Qwen3.5-9B at 82.5 on MMLU-Pro and 91.5 on IFEval, against 80.8 and 88.9 for OpenAI’s gpt-oss-120b.
- Google reports Gemma 4 E4B at 69.4% on MMLU Pro and 42.2% on Tau2, against 67.6% and 16.2% for the previous generation’s Gemma 3 27B.
- Mistral describes Ministral 3 14B as comparable to its larger Mistral Small 3.2 24B.
Large models still win on hard code generation, competition maths and broad knowledge. In the same Qwen table, gpt-oss-120b leads the 9B on LiveCodeBench v6 by 82.7 to 65.6. Liquid AI does not recommend LFM2.5-2.6B for agentic coding or knowledge-heavy tasks. A position paper by Belcak and colleagues, revised in September 2026, argues that small models are capable, better suited and cheaper for many of the individual calls inside agent systems, and that agents should mix models rather than send everything to one.
A practical way to decide:
- List the tasks, not the use case. One agent may classify, extract, call tools and write a summary.
- Run each task through a 4B and a 9B model on your own test set before you consider anything larger.
- Escalate only the tasks that fail: broad knowledge, long multi-step code or open-ended reasoning.
- Fine-tune where volume justifies it. A small model trained on your examples often closes the remaining gap on a narrow task.
- Route at runtime, so each request reaches the smallest model that passed. Our article on SLMs in enterprise infrastructure covers the operating model.
Licences to read before you ship
Most of this list is permissive. Qwen3.5, Gemma 4, Ministral 3, Granite 4.2, SmolLM3, MiniCPM5 and Olmo 3 use Apache 2.0, and Microsoft’s Phi-4 models use MIT. Ai2 adds that Olmo 3 is intended for research and educational use under its responsible-use guidelines. Three licences need a closer read:
- Llama 3.2 Community License. Companies whose products had more than 700 million monthly active users when the version was released must request a separate licence from Meta, and redistributors must display “Built with Llama”. Meta’s EU restriction applies to its multimodal models, not to the text-only 1B and 3B.
- LFM Open License 1.0. Liquid AI grants commercial rights only to organisations below $10 million in annual revenue, with an exception for qualifying non-profits doing research. Most enterprises need a separate agreement.
- NVIDIA Nemotron Open Model License. NVIDIA states that Nemotron 3 Nano 4B is ready for commercial use; have counsel read the terms as you would any custom licence.
Our guide to open-weight licence terms covers the review process.
How VDF AI fits
VDF AI Router is built around the mix this guide recommends: small tasks go to small models, and larger models answer only where they earn it. Local Ollama and custom on-premises deployments register alongside any cloud models your policy allows, and allow and deny lists, regulated-domain approvals and an air-gapped local-only mode are applied before any learned routing decision.
For narrow, high-volume tasks, VDF’s fine-tuning workflow lets teams choose a base model, build a dataset from governed sources and train it on their own infrastructure, with evaluation gates before promotion. The Model Evaluation Suite runs your test prompts against registered models, local Ollama deployments included, inside your deployment, so the choice between a 4B and a 9B model rests on your results.
Sources
Verified 6 October 2026.
- Qwen3.5-9B model card and Qwen3.5-0.8B model card
- Qwen3.5-4B and Qwen3.5-2B model cards
- Gemma 4 E4B model card and Gemma 4 12B model card
- Ministral 3 8B, Ministral 3 14B and Ministral 3 3B model cards
- Granite 4.2 8B model card and Granite 4.2 3B
- Phi-4-mini-instruct model card and Phi-4 model card
- Llama 3.2 model card, licence and acceptable use policy
- SmolLM3-3B model card
- LFM2.5-2.6B, LFM2.5-8B-A1B and LFM Open License 1.0
- NVIDIA Nemotron 3 Nano 4B model card
- MiniCPM5-2B model card and MiniCPM-V 4.6 model card
- Olmo 3 7B Instruct model card
- gpt-oss-20b model card
- Berkeley Function Calling Leaderboard
- Belcak et al., Small Language Models are the Future of Agentic AI
- Llama 3.2 3B config.json (public copy of the gated file) and SmolLM3-3B config.json
- llama.cpp quantize README: Q4_K_M bits per weight
Mixing small and large models on your own hardware? See how VDF AI Router routes between them, or book a demo.