Small Language Models

Best Small Language Models (2026): SLMs Compared by Task

Eleven open-weight small language model families from 0.8B to 14B parameters, checked against their model cards in October 2026 and compared by task: tool calling, retrieval over long documents, image input and edge deployment, with licences, memory arithmetic and when a small model beats a large one.

The best small language models in 2026 are Qwen3.5-9B and Granite 4.2 8B for tool calling and agents, Gemma 4 E4B and Ministral 3 8B for image input on modest hardware, and Phi-4-mini or Llama 3.2 3B where only a 3B footprint fits. All except Llama 3.2 are Apache 2.0 or MIT, and each fits an 8–16 GB GPU at 4-bit.

Small language models changed more in the last year than their parameter counts suggest. Current releases between 1 and 14 billion parameters handle tool calls, long documents and images, and several publish agent benchmarks on their model cards. For a platform team, that changes the default question from “which large model” to “which tasks still need one”.

This guide compares eleven open-weight families. Licences, parameter counts, context lengths and tool-calling notes were checked against the vendors’ model cards on Hugging Face in October 2026. Scores appear only where a vendor or a public leaderboard published them, with the benchmark named. For the definition and the basic trade-offs, see our small language model explainer.

Quick picks by task

TaskPickWhy, from the model card
Tool calling and agentsQwen3.5-9BQwen reports 66.1 on BFCL-V4 and 79.1 on TAU2-Bench; vLLM tool parser documented
Tool calling with a dense, text-only modelGranite 4.2 8BIBM lists tool calling and agentic workflows as primary uses; 128K context, extendable to 512K
Retrieval over long documentsQwen3.5-4B or Qwen3.5-9B262,144-token native context with a KV cache in only 8 of 32 layers
Text, images and audio on a laptopGemma 4 E4BGoogle built the E models for laptops and mobile devices
Images plus function callingMinistral 3 8B0.4B vision encoder, native function calling and JSON output; fits 12 GB in FP8, per Mistral
The strongest model under 15BGemma 4 12B or Ministral 3 14BGoogle reports 77.2% on MMLU Pro for the 12B; Mistral compares the 14B with its 24B Mistral Small 3.2
Phones and CPUsQwen3.5-0.8B, Llama 3.2 1B or LFM2.5-1.2BUnder 1 GB of weights at 4-bit
Fully open training dataSmolLM3-3B or Olmo 3 7BWeights, data mixtures and training details published

What counts as a small language model in 2026

The working range today is roughly 0.5 to 14 billion parameters, small enough to run on one consumer GPU, a laptop or, at the bottom end, a phone. Two architectures stretch the definition. Mixture-of-experts models such as Liquid AI’s LFM2.5-8B-A1B keep 8.3 billion parameters in memory but activate 1.5 billion per token. Google’s Gemma 4 E2B and E4B count “effective” parameters: 2.3B and 4.5B, or 5.1B and 8B including their per-layer embeddings.

Three things separate the 2026 generation from the 7B models of 2024:

  1. Built-in vision. Every Qwen3.5 size includes a vision encoder, Ministral 3 adds a 0.4B one, and Gemma 4’s small models accept audio as well as images.
  2. Long context at low cost. Qwen3.5 and LFM2.5 mix full attention with linear or convolution layers, so only some layers keep a growing KV cache. A 262K-token window on a 9B model is now practical on one GPU.
  3. Training for tool use. Model cards now report function-calling and agent benchmarks such as BFCL and TAU2, and name the vLLM parser that reads the model’s tool calls.

The shortlist: eleven small model families

Checked against the Hugging Face model cards in October 2026. Parameter counts are the Hub totals, which include vision or audio encoders where a model has them.

Model (vendor)YearParametersContextInputsLicenceTool calling, per the card
Qwen3.5-0.8B, 2B, 4B, 9B (Qwen)20260.9B, 2.3B, 4.7B, 9.7B262KText, image, videoApache 2.0Yes; vLLM parser qwen3_coder
Gemma 4 E2B, E4B (Google)20262.3B and 4.5B effective128KText, image, audioApache 2.0Native function calling
Gemma 4 12B (Google)202612.0B256KText, image, audioApache 2.0Native function calling
Ministral 3 3B, 8B, 14B (Mistral)20253.8B, 8.9B, 13.9B256KText, imageApache 2.0Native function calling and JSON; parser mistral
Granite 4.2 3B, 8B (IBM)20263.7B, 8.8B128K, extendable to 512KTextApache 2.0Reasoning before tool calls; parser qwen3_coder
Phi-4-mini-instruct (Microsoft)20253.8B128KTextMITDedicated function-calling format
Llama 3.2 1B, 3B (Meta)20241.2B, 3.2B128KTextLlama 3.2 Community LicenseAimed at agentic retrieval and summarization
SmolLM3-3B (Hugging Face)20253.1B64K, 128K with YaRNTextApache 2.0XML or Python-style calls; parser hermes
LFM2.5-1.2B, 2.6B, 8B-A1B (Liquid AI)20261.2B, 2.7B, 8.3B with 1.5B active32K for 1.2B; 128K for the othersTextLFM Open License 1.0Pythonic function calls
Nemotron 3 Nano 4B (NVIDIA)20264.0BUp to 262KText, English onlyNVIDIA Nemotron Open Model LicenseYes; parser qwen3_coder
MiniCPM5-2B (OpenBMB)20262.5B128KTextApache 2.0Designed for tool-use workflows
Olmo 3 7B Instruct (Ai2)20257.3B64KText, EnglishApache 2.0Function-calling system prompt

A few notes behind the table. Granite 4.2’s card gives 25 August 2026 as the release date, and NVIDIA’s gives 16 March 2026 for Nemotron 3 Nano 4B, which NVIDIA compressed from its 9B Nemotron Nano v2. Qwen lists the 0.8B model’s intended uses as prototyping, task-specific fine-tuning and research. Just above the range, OpenAI’s gpt-oss-20b has 21B parameters with 3.6B active and runs within 16 GB of memory.

Best small language model for tool calling

Vendors now publish function-calling scores, but each runs its own harness and settings. Compare sizes within one family, not numbers across families.

FamilyBenchmarkSmallest → largest size in the family
Qwen3.5BFCL-V40.8B: 25.3 · 2B: 43.6 · 4B: 50.3 · 9B: 66.1
Qwen3.5TAU2-Bench0.8B: 11.6 · 2B: 48.8 · 4B: 79.9 · 9B: 79.1
Granite 4.2BFCL v43B: 52.41 · 8B: 52.39
Granite 4.2τ³-bench3B: 50.99 · 8B: 66.34
Gemma 4Tau2, average of threeE2B: 24.5% · E4B: 42.2% · 12B: 69.0%

Two patterns stand out. In every family that publishes these numbers, scores climb steeply from the sub-1B and 2B sizes to 4B. And Qwen3.5-4B matches the 9B on TAU2-Bench while trailing it clearly on BFCL-V4, so test both when memory is tight.

The independent Berkeley Function Calling Leaderboard, last updated on 12 April 2026, does not yet list the 2026 small models, but its older entries show the same direction. Qwen3-8B in function-calling mode scored 42.57% overall, ahead of Mistral Large 2411 at 38.37%. Llama 3.2 3B scored 21.95% and Llama 3.2 1B 10.82%. Salesforce’s xLAM-2-8b-fc-r reached 46.68%, but its CC-BY-NC-4.0 licence rules out commercial use.

Three settings from the model cards make small models more reliable with tools:

  • Use the parser the card names. vLLM needs --enable-auto-tool-choice plus the right parser: qwen3_coder for Qwen3.5, Granite 4.2 and Nemotron 3 Nano 4B, mistral for Ministral 3 and hermes for SmolLM3.
  • Keep the tool list short. Mistral advises limiting tools to the minimum the use case needs, and a temperature below 0.1 in production.
  • Follow the sampling advice. NVIDIA recommends a temperature of 0.6 for tool calling with Nemotron 3 Nano 4B, lower than for reasoning tasks.

When a small model is not enough, our agent and tool-calling picks cover the larger models, and the VRAM tier guide shows what each card can hold.

RAG and long documents: context costs memory

Retrieval pipelines put long passages into every prompt, so the KV cache, not the weights, often decides whether a small model fits. Per token at 16 bits it equals 2 × layers with a cache × KV heads × head dimension × 2 bytes, from each model’s config.json:

ModelLayers with a growing KV cacheKV cache per tokenOne 32K-token conversationOne 128K-token conversation
Qwen3.5-0.8B6 of 2412 KiB0.4 GB1.6 GB
Qwen3.5-9B8 of 3232 KiB1.1 GB4.3 GB
Llama 3.2 1B16 of 1632 KiB1.1 GB4.3 GB
SmolLM3-3B36 of 3672 KiB2.4 GB9.7 GB
Llama 3.2 3B28 of 28112 KiB3.8 GB15.0 GB
Phi-4-mini32 of 32128 KiB4.3 GB17.2 GB
Ministral 3 8B34 of 34136 KiB4.6 GB18.3 GB
Granite 4.2 8B40 of 40160 KiB5.4 GB21.5 GB

The difference is architectural. Granite 4.2 8B is a dense transformer that caches every layer, so at 128K tokens its cache is four times the size of its 4-bit weights. Qwen3.5-9B caches only its 8 full-attention layers; the other 24 carry a fixed-size state instead. For one GPU serving several long-context retrieval sessions, the hybrid design is the easier fit. Liquid AI likewise recommends LFM2.5-2.6B for retrieval and data extraction, but not for knowledge-heavy tasks.

Two caveats. A long context window does not mean a model uses all of it well, so test retrieval quality at the lengths you plan to serve. And retrieval usually beats stuffing whole documents into the prompt; our note on long context versus retrieval explains when each wins.

Hardware and edge deployment

Weights at three precisions, using the same arithmetic as our VRAM calculator: 16 bits, 8 bits and 4.89 bits per weight for GGUF Q4_K_M.

Size classExamplesBF16FP8Q4_K_MTypical hardware
Under 1.5BQwen3.5-0.8B, Llama 3.2 1B, LFM2.5-1.2B1.7–2.5 GB0.9–1.2 GB0.5–0.8 GBPhone, CPU or any GPU
2–4BQwen3.5-4B, Phi-4-mini, Ministral 3 3B, Granite 4.2 3B, SmolLM3-3B6.2–9.3 GB3.1–4.7 GB1.9–2.8 GBLaptop or an 8 GB GPU
7–9BQwen3.5-9B, Granite 4.2 8B, Ministral 3 8B, Olmo 3 7B14.6–19.3 GB7.3–9.7 GB4.5–5.9 GBA 12–16 GB GPU
12–14BGemma 4 12B, Ministral 3 14B, Phi-423.9–29.3 GB12.0–14.7 GB7.3–9.0 GBA 16–24 GB GPU

Add the KV cache from the previous table for each conversation you serve at once, plus about a tenth of memory for the runtime. A 16 GB card holds a 9B model at 4-bit with several long conversations; our GPU buyer’s guide compares the cards by budget.

Small models with vision, audio and edge targets

Several small models now read images, which removes a separate vision service from many pipelines.

  • Qwen3.5 accepts image and video input at every size from 0.8B up. Qwen’s tables show the 9B ahead of its own previous Qwen3-VL-30B-A3B on most vision benchmarks it reports.
  • Gemma 4 E2B and E4B take text, images and audio, use 512-token sliding-window layers to save memory, and let you trade detail for speed with visual token budgets from 70 to 1,120 tokens per image.
  • Ministral 3 pairs each model with a 0.4B vision encoder. Mistral says the 3B fits in 8 GB of VRAM at FP8, the 8B in 12 GB and the 14B in 24 GB.
  • MiniCPM-V 4.6 builds a 1.3B vision model on Qwen3.5-0.8B, and OpenBMB demonstrates it on iOS, Android and HarmonyOS phones.
  • Nemotron 3 Nano 4B targets edge platforms such as Jetson Thor, GeForce RTX cards and DGX Spark, in English only.

Our guide to the best local vision models compares the larger vision models as well.

Small language models vs large language models

Small models win on narrow, repeated work: classification, extraction, routing, short answers grounded in retrieved text and well-defined tool calls. They answer faster, fit on hardware you already own and can be fine-tuned on your own examples. The published evidence for how close they come:

  • Qwen’s own table puts Qwen3.5-9B at 82.5 on MMLU-Pro and 91.5 on IFEval, against 80.8 and 88.9 for OpenAI’s gpt-oss-120b.
  • Google reports Gemma 4 E4B at 69.4% on MMLU Pro and 42.2% on Tau2, against 67.6% and 16.2% for the previous generation’s Gemma 3 27B.
  • Mistral describes Ministral 3 14B as comparable to its larger Mistral Small 3.2 24B.

Large models still win on hard code generation, competition maths and broad knowledge. In the same Qwen table, gpt-oss-120b leads the 9B on LiveCodeBench v6 by 82.7 to 65.6. Liquid AI does not recommend LFM2.5-2.6B for agentic coding or knowledge-heavy tasks. A position paper by Belcak and colleagues, revised in September 2026, argues that small models are capable, better suited and cheaper for many of the individual calls inside agent systems, and that agents should mix models rather than send everything to one.

A practical way to decide:

  1. List the tasks, not the use case. One agent may classify, extract, call tools and write a summary.
  2. Run each task through a 4B and a 9B model on your own test set before you consider anything larger.
  3. Escalate only the tasks that fail: broad knowledge, long multi-step code or open-ended reasoning.
  4. Fine-tune where volume justifies it. A small model trained on your examples often closes the remaining gap on a narrow task.
  5. Route at runtime, so each request reaches the smallest model that passed. Our article on SLMs in enterprise infrastructure covers the operating model.

Licences to read before you ship

Most of this list is permissive. Qwen3.5, Gemma 4, Ministral 3, Granite 4.2, SmolLM3, MiniCPM5 and Olmo 3 use Apache 2.0, and Microsoft’s Phi-4 models use MIT. Ai2 adds that Olmo 3 is intended for research and educational use under its responsible-use guidelines. Three licences need a closer read:

  • Llama 3.2 Community License. Companies whose products had more than 700 million monthly active users when the version was released must request a separate licence from Meta, and redistributors must display “Built with Llama”. Meta’s EU restriction applies to its multimodal models, not to the text-only 1B and 3B.
  • LFM Open License 1.0. Liquid AI grants commercial rights only to organisations below $10 million in annual revenue, with an exception for qualifying non-profits doing research. Most enterprises need a separate agreement.
  • NVIDIA Nemotron Open Model License. NVIDIA states that Nemotron 3 Nano 4B is ready for commercial use; have counsel read the terms as you would any custom licence.

Our guide to open-weight licence terms covers the review process.

How VDF AI fits

VDF AI Router is built around the mix this guide recommends: small tasks go to small models, and larger models answer only where they earn it. Local Ollama and custom on-premises deployments register alongside any cloud models your policy allows, and allow and deny lists, regulated-domain approvals and an air-gapped local-only mode are applied before any learned routing decision.

For narrow, high-volume tasks, VDF’s fine-tuning workflow lets teams choose a base model, build a dataset from governed sources and train it on their own infrastructure, with evaluation gates before promotion. The Model Evaluation Suite runs your test prompts against registered models, local Ollama deployments included, inside your deployment, so the choice between a 4B and a 9B model rests on your results.

Sources

Verified 6 October 2026.


Mixing small and large models on your own hardware? See how VDF AI Router routes between them, or book a demo.

Frequently asked questions

What is the best small language model in 2026?

There is no single winner, because the right model depends on the task. For tool calling and agents, Qwen3.5-9B posts the highest function-calling scores in its family and documents its tool parser, and Granite 4.2 8B is a text-only alternative from IBM. For image input on modest hardware, Gemma 4 E4B and Ministral 3 8B stand out. Where only about 3B fits, Phi-4-mini and Llama 3.2 3B are long-established choices. Most use Apache 2.0 or MIT licences, so test two candidates on your own tasks.

What is the best small language model for tool calling?

Among models under 10B parameters, Qwen3.5-9B is the first one to test. Qwen reports 66.1 on the BFCL-V4 function-calling benchmark and 79.1 on TAU2-Bench, and the model card documents the vLLM tool parser to use. IBM's Granite 4.2 8B lists tool calling and agentic workflows among its primary uses, and Ministral 3 8B adds image input with native function calling. Each vendor runs its own harness, so confirm the ranking with your own tools and prompts before choosing.

Are small language models better than large language models?

For narrow, repeated tasks they often match larger models at a fraction of the memory and latency. Qwen reports its 9B model ahead of gpt-oss-120b on MMLU-Pro and instruction following, and Google reports Gemma 4 E4B ahead of the older Gemma 3 27B. Large models still lead on hard coding, competition maths and broad world knowledge; the same Qwen table shows gpt-oss-120b well ahead on LiveCodeBench. Most enterprises run both and route each request to the smallest model that passes their tests.

What hardware does a small language model need?

The weights are modest: about 2 GB for a 3B model and 6 GB for a 9B model at 4-bit, or about 1.6 times that at 8-bit. Context is the variable to watch. Granite 4.2 8B stores 160 KiB of KV cache per token at 16 bits, so one 128,000-token conversation needs about 21 GB on top of the weights. Qwen3.5-9B keeps a cache in only a quarter of its layers and needs about 4.3 GB for the same length. An 8 to 16 GB GPU covers most small-model deployments.

Can small language models run on a CPU or a phone?

The smallest ones can. Models under about 1.5B parameters, such as Qwen3.5-0.8B, Llama 3.2 1B and LFM2.5-1.2B, need well under 2 GB of memory at 4-bit, and their vendors target on-device use. Meta's Llama 3.2 card describes quantized versions for devices with limited compute, and OpenBMB demonstrates MiniCPM-V 4.6 on iOS, Android and HarmonyOS phones. Expect slower generation than on a GPU, and test whether a model this small is accurate enough for your task.

Filed under
small language modelsopen-weight modelstool callinglocal LLMon-premises AImodel routing
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading