The best local LLM for agents and tool calling emits well-formed calls your server can parse, picks the right tool across many turns, and keeps its context affordable over a long loop. In October 2026 that means Qwen3.6-35B-A3B or GLM-4.7-Flash on a 24 to 32 GB GPU, Qwen3.5-9B on 8 to 16 GB, and gpt-oss-120b or Mistral Small 4 on 128 GB machines.
An agent is a model in a loop: it reads the task, calls a tool, reads the result and decides what to do next, often dozens of times. A model that writes good prose can still fail that loop by inventing a parameter, picking the wrong tool, or running out of context on turn twenty. This guide picks models on the behaviours that matter for agents and shows how to serve them so their tool calls arrive intact.
Model facts and vendor benchmarks were checked on each Hugging Face card, and serving behaviour on the vLLM, llama.cpp and Ollama documentation, on 6 October 2026. The design patterns around the model, such as approval gates and idempotent writes, are covered in our tool calling patterns.
Quick picks
| Hardware | Pick | Evidence, as reported by the vendor | vLLM tool parser |
|---|---|---|---|
| 8–16 GB GPU | Qwen3.5-9B | 66.1 on BFCL-V4, 79.1 on TAU2-Bench (Qwen) | qwen3_coder |
| 16 GB GPU | gpt-oss-20b | Native function calling and structured outputs (OpenAI) | openai |
| 24–32 GB GPU | Qwen3.6-35B-A3B | 62.8 on MCP-Atlas, 67.2 on TAU3-Bench (Qwen) | qwen3_coder |
| 24–32 GB GPU, MIT licence | GLM-4.7-Flash | 79.5 on τ²-Bench (Z.ai) | glm47 |
| 64–128 GB GPU or unified memory | gpt-oss-120b, Mistral Small 4 or Qwen3.5-122B-A10B | Qwen3.5-122B-A10B: 72.2 on BFCL-V4, 79.5 on TAU2-Bench (Qwen) | openai, mistral, qwen3_coder |
What makes a model good at tool calling
Five behaviours separate a dependable agent model from one that demos well:
- Format adherence. Every call matches the format the model was trained on, so the server can turn it into a structured
tool_callsfield. - Tool selection and restraint. It picks the right tool among many and calls none when no tool fits. The Berkeley Function Calling Leaderboard scores the second skill separately as irrelevance detection.
- Argument accuracy. Parameter names, types and required fields match the JSON schema you supplied.
- Multi-turn state. It carries results from earlier calls into later ones and stops when the task is done.
- Recovery. When a tool returns an error, it reads the message and retries sensibly instead of repeating the same call.
Single-call tests measure the first three. Agent work depends on the last two.
The shortlist
| Model | Total / active parameters | 4-bit file | Native window | Licence | Agent evidence (vendor-reported) | vLLM parser |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | 9.7B dense | 5.6 GB | 262K | Apache 2.0 | BFCL-V4 66.1; TAU2-Bench 79.1 | qwen3_coder |
| Qwen3.6-35B-A3B | 36B / 3B | 20.4 GB | 262K | Apache 2.0 | MCP-Atlas 62.8; MCPMark 37.0; TAU3-Bench 67.2 | qwen3_coder |
| Qwen3.8-27B | 27.8B dense | 19.0 GB | 262K | Apache 2.0 | CoWorkBench 70.7; IFBench 79.5 | See Qwen’s vLLM recipe |
| GLM-4.7-Flash | 31.2B / about 3B | 18.2 GB | 202,752 (config) | MIT | τ²-Bench 79.5 | glm47 |
| Gemma 4 31B | 31.3B dense | 18.0 GB | 256K | Apache 2.0 | Tau2 76.9% (reported as an average over 3) | Native function calling per Google |
| Gemma 4 26B A4B | 25.8B / 3.8B | 14.6 GB | 256K | Apache 2.0 | Tau2 68.2% | Native function calling per Google |
| gpt-oss-20b / gpt-oss-120b | 21B / 3.6B; 117B / 5.1B | 12.1 GB; 63.4 GB | 131,072 | Apache 2.0 plus usage policy | Function calling, browsing, Python tool, structured outputs | openai |
| Granite 4.2 8B / 30B | 8.8B; 29.3B dense | 5.4 GB; 17.7 GB | 128K | Apache 2.0 | BFCL v4 52.39 / 61.39; τ³-bench 66.34 / 68.05 | qwen3_coder per IBM |
| Ministral 3 14B | 13.9B | 8.2 GB | 256K | Apache 2.0 | Native function calling and JSON output | mistral |
| Mistral Small 4 | 119B / 6.5B | 72.2 GB GGUF; 67.8 GB MLX | 256K | Apache 2.0 | Native function calling and JSON output | mistral |
File sizes are the published GGUF builds (ggml-org, or the vendor’s own for Ministral and Granite, LM Studio’s for Qwen3.5-9B and Mistral Small 4); gpt-oss ships natively in MXFP4. Our VRAM tier guide and Mac guide show how much context each leaves on common hardware.
Two models were left off. MiniMax M2.7 is licensed for non-commercial use only. Qwen-AgentWorld-35B-A3B is a world model that simulates agent environments, not a general agent model.
Reading agent benchmarks without being misled
Vendor harnesses differ. Qwen evaluates TAU2-Bench with the airline-domain fixes from Anthropic’s Claude Opus 4.5 system card. Z.ai applied the same airline fixes and also added a prompt to the retail and telecom user simulations. Google reports Tau2 as an average over three without detailing its setup on the card. Each choice is reasonable, and together they mean a 79.1, a 79.5 and a 76.9 cannot be put in rank order.
Benchmark versions move. The τ-bench repository now ships τ³-bench, which adds a knowledge-retrieval banking domain, voice evaluation and more than 75 task fixes. Its July 2026 grading update states that results from versions before 1.0.1 are not comparable with later ones.
Single calls overstate agent skill. The Berkeley Function Calling Leaderboard, last updated 12 April 2026, scores single, parallel and multi-turn calling separately. The gap is large for most open models:
| Model (BFCL, April 2026) | Live single-call accuracy | Multi-turn accuracy |
|---|---|---|
| Qwen3-32B (FC) | 82.01% | 47.87% |
| GLM-4.6 (FC, thinking) | 80.90% | 68.00% |
| Llama 3.3 70B Instruct (FC) | 76.61% | 21.50% |
| Gemma 3 27B (prompt) | 74.54% | 10.75% |
The board predates Qwen3.5, Gemma 4 and gpt-oss, so use it for the pattern rather than a ranking of current models.
One harness for several models. NVIDIA’s Nemotron 3.5 Lightning card runs Qwen3.6-35B-A3B, Gemma 4 26B A4B, gpt-oss-20b and its own models through one harness. Qwen3.6-35B-A3B led most agentic rows there, including 48.74 on BrowseComp against 26.30 for Gemma 4 26B A4B, while Gemma 4 26B A4B led instruction following with 77.25 on IFBench.
Serving: parsers, strict mode and structured output
A model’s tool-call format only becomes a structured tool_calls response if the server parses it.
vLLM. Start the server with --enable-auto-tool-choice and the --tool-call-parser the model card names. vLLM’s tool calling documentation lists parsers including hermes, mistral, llama3_json, llama4_pythonic, openai for gpt-oss, glm45 and glm47, deepseek_v31, kimi_k2 and cohere_command4; the Qwen3.5 and Qwen3.6 cards specify qwen3_coder. Three settings decide how much you can trust the output:
tool_choice="required"and named function calls use structured outputs, so the response is guaranteed to parse against your schema.- With
tool_choice="auto", vLLM constrains a tool’s arguments only when the tool setsstrict: true, or when the operator raises--tool-strict-level. Its documentation notes that most OpenAI-compatible clients never setstrict. - Otherwise, the docs warn, arguments may occasionally be malformed or violate the function’s schema.
llama.cpp. llama-server supports OpenAI-style function calling through Jinja chat templates, which current builds enable by default. Native handlers cover formats such as Llama 3.x, Hermes 2 and 3, Qwen 2.5, Mistral Nemo and GPT-OSS harmony; other templates fall back to a generic handler that, per the function-calling guide, may use more tokens. Parallel calls are off unless the request sets "parallel_tool_calls": true, and response_format accepts a JSON schema.
Ollama. The chat API takes a tools array and documents single, parallel and multi-turn agent loops. Structured outputs take a JSON schema in the format field, and Ollama advises repeating the schema in the prompt. Recent releases fixed Gemma 4 tool-response continuations and GLM tool calls dropped at the end of generation, so stay current.
For a wider comparison of these servers, see our vLLM, Ollama and llama.cpp guide.
Context budget for tool loops
Tool definitions, calls, results and reasoning all accumulate in the context, and every agent run active at once needs its own KV cache. Ollama’s documentation recommends at least 64,000 tokens for agents, and Qwen advises keeping at least 128K for Qwen3.6 to preserve thinking.
Worked example: parallel 64K agent runs on a 32 GB GPU. Using the method of our sizing calculator, 28.8 GB is usable. Subtract the weights file and divide by the cache one 64K-token run needs:
| Model | Weights | Cache per 64K run | Runs that fit (16-bit) | With an 8-bit cache |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | 20.4 GB | 1.41 GB | 5 | 11 |
| gpt-oss-20b | 12.1 GB | 1.61 GB | 10 | 20 |
| GLM-4.7-Flash | 18.2 GB | 3.55 GB | 2 | 5 |
| Qwen3.8-27B | 19.0 GB | 4.45 GB | 2 | 4 |
| Gemma 4 31B | 18.0 GB | 6.21 GB | 1 | 3 |
Qwen3.6-35B-A3B keeps a growing cache in only 10 of 40 layers, which is why it supports five runs where a dense 31B supports one. On 128 GB of unified memory, gpt-oss-120b fits 16 such runs, Qwen3.5-122B-A10B 18 and Mistral Small 4 22. The general sizing method is in our GPU requirements guide.
Two habits stretch the budget. Expose only the tools a task needs, because every schema is resent on every turn. And trim old tool output: Qwen’s own search agents drop earlier tool responses once their combined length passes a threshold, a strategy its cards call context folding.
Reasoning traces inside tool loops
Most current models think before they act, and each family wants its reasoning handled differently across turns:
- Gemma 4. Google’s card says earlier thoughts must not be added to the history before the next user turn, except on tool-call turns, where thinking content should be preserved.
- Qwen3.8. Preserved thinking is on by default, and
reasoning_effortaccepts low, medium and xhigh. - GLM-4.7-Flash. Z.ai turns on Preserved Thinking mode for its multi-turn agent benchmarks and serves it with the
glm45reasoning parser. - gpt-oss. OpenAI says the models work correctly only with its harmony response format, and reasoning effort can be set to low, medium or high.
Get this wrong and the model loses track of why it called a tool, or the context fills with stale reasoning. Use the chat template shipped with the checkpoint and the reasoning parser its card names, and test a full multi-turn loop rather than a single call. If you run a desktop agent harness rather than your own loop, our OpenClaw and Hermes Agent setup guide covers that configuration. For agents that mainly retrieve documents, see our local RAG model picks. If the agent has to run on a laptop or a 16 GB card, start from the small models ranked for tool calling.
How VDF AI fits
VDF AI Agents treats the model as one setting of an agent. Its product page lists connecting any LLM or small model per agent, including Ollama and other compatible endpoints, and assigning tools from an MCP-powered registry. Tool access is scoped per agent, high-risk actions can require human approval, and the audit log records every prompt, tool call and output.
The MCP gateway governs the tools themselves: administrators grant tools to roles in a registry, the MCP server layer runs inside your perimeter, including air-gapped estates, and each tool call leaves an audit record. That combination lets you swap the local model behind an agent, for example from Qwen3.6-35B-A3B to GLM-4.7-Flash, without changing which tools the agent may touch.
Sources
Model cards, leaderboards and documentation verified 6 October 2026.
- Model cards: Qwen3.5-9B, Qwen3.5-122B-A10B, Qwen3.6-35B-A3B, Qwen3.8-27B, GLM-4.7-Flash, Gemma 4 31B, gpt-oss-20b, gpt-oss-120b, Granite 4.2 8B, Ministral 3 14B, Mistral Small 4, Nemotron 3.5 Lightning, Qwen-AgentWorld-35B-A3B
- MiniMax M2.7 licence
- Berkeley Function Calling Leaderboard (updated 12 April 2026)
- τ-bench and τ³-bench repository
- vLLM tool calling and vLLM structured outputs
- llama.cpp function calling and llama.cpp server README
- Ollama tool calling, Ollama structured outputs, Ollama context length and Ollama releases
- GGUF file sizes: ggml-org Qwen3.6-35B-A3B, ggml-org GLM-4.7-Flash, ggml-org gpt-oss-120b, IBM Granite 4.2 30B GGUF, LM Studio Mistral Small 4 GGUF
Putting local models behind governed agents? See how VDF AI Agents scopes tools per agent, or book a demo.