AI Infrastructure

Best Local LLM for Agents and Tool Calling (2026)

Which open-weight model to run behind a local AI agent in October 2026, judged on what agent loops need: well-formed function calls, the right tool across many turns, structured output, a context budget that survives long tool loops, and a serving stack that parses the model's tool-call format. Vendor benchmarks with their caveats, memory sizing for parallel agent runs, and parser settings for vLLM, llama.cpp and Ollama.

The best local LLM for agents and tool calling emits well-formed calls your server can parse, picks the right tool across many turns, and keeps its context affordable over a long loop. In October 2026 that means Qwen3.6-35B-A3B or GLM-4.7-Flash on a 24 to 32 GB GPU, Qwen3.5-9B on 8 to 16 GB, and gpt-oss-120b or Mistral Small 4 on 128 GB machines.

An agent is a model in a loop: it reads the task, calls a tool, reads the result and decides what to do next, often dozens of times. A model that writes good prose can still fail that loop by inventing a parameter, picking the wrong tool, or running out of context on turn twenty. This guide picks models on the behaviours that matter for agents and shows how to serve them so their tool calls arrive intact.

Model facts and vendor benchmarks were checked on each Hugging Face card, and serving behaviour on the vLLM, llama.cpp and Ollama documentation, on 6 October 2026. The design patterns around the model, such as approval gates and idempotent writes, are covered in our tool calling patterns.

Quick picks

HardwarePickEvidence, as reported by the vendorvLLM tool parser
8–16 GB GPUQwen3.5-9B66.1 on BFCL-V4, 79.1 on TAU2-Bench (Qwen)qwen3_coder
16 GB GPUgpt-oss-20bNative function calling and structured outputs (OpenAI)openai
24–32 GB GPUQwen3.6-35B-A3B62.8 on MCP-Atlas, 67.2 on TAU3-Bench (Qwen)qwen3_coder
24–32 GB GPU, MIT licenceGLM-4.7-Flash79.5 on τ²-Bench (Z.ai)glm47
64–128 GB GPU or unified memorygpt-oss-120b, Mistral Small 4 or Qwen3.5-122B-A10BQwen3.5-122B-A10B: 72.2 on BFCL-V4, 79.5 on TAU2-Bench (Qwen)openai, mistral, qwen3_coder

What makes a model good at tool calling

Five behaviours separate a dependable agent model from one that demos well:

  1. Format adherence. Every call matches the format the model was trained on, so the server can turn it into a structured tool_calls field.
  2. Tool selection and restraint. It picks the right tool among many and calls none when no tool fits. The Berkeley Function Calling Leaderboard scores the second skill separately as irrelevance detection.
  3. Argument accuracy. Parameter names, types and required fields match the JSON schema you supplied.
  4. Multi-turn state. It carries results from earlier calls into later ones and stops when the task is done.
  5. Recovery. When a tool returns an error, it reads the message and retries sensibly instead of repeating the same call.

Single-call tests measure the first three. Agent work depends on the last two.

The shortlist

ModelTotal / active parameters4-bit fileNative windowLicenceAgent evidence (vendor-reported)vLLM parser
Qwen3.5-9B9.7B dense5.6 GB262KApache 2.0BFCL-V4 66.1; TAU2-Bench 79.1qwen3_coder
Qwen3.6-35B-A3B36B / 3B20.4 GB262KApache 2.0MCP-Atlas 62.8; MCPMark 37.0; TAU3-Bench 67.2qwen3_coder
Qwen3.8-27B27.8B dense19.0 GB262KApache 2.0CoWorkBench 70.7; IFBench 79.5See Qwen’s vLLM recipe
GLM-4.7-Flash31.2B / about 3B18.2 GB202,752 (config)MITτ²-Bench 79.5glm47
Gemma 4 31B31.3B dense18.0 GB256KApache 2.0Tau2 76.9% (reported as an average over 3)Native function calling per Google
Gemma 4 26B A4B25.8B / 3.8B14.6 GB256KApache 2.0Tau2 68.2%Native function calling per Google
gpt-oss-20b / gpt-oss-120b21B / 3.6B; 117B / 5.1B12.1 GB; 63.4 GB131,072Apache 2.0 plus usage policyFunction calling, browsing, Python tool, structured outputsopenai
Granite 4.2 8B / 30B8.8B; 29.3B dense5.4 GB; 17.7 GB128KApache 2.0BFCL v4 52.39 / 61.39; τ³-bench 66.34 / 68.05qwen3_coder per IBM
Ministral 3 14B13.9B8.2 GB256KApache 2.0Native function calling and JSON outputmistral
Mistral Small 4119B / 6.5B72.2 GB GGUF; 67.8 GB MLX256KApache 2.0Native function calling and JSON outputmistral

File sizes are the published GGUF builds (ggml-org, or the vendor’s own for Ministral and Granite, LM Studio’s for Qwen3.5-9B and Mistral Small 4); gpt-oss ships natively in MXFP4. Our VRAM tier guide and Mac guide show how much context each leaves on common hardware.

Two models were left off. MiniMax M2.7 is licensed for non-commercial use only. Qwen-AgentWorld-35B-A3B is a world model that simulates agent environments, not a general agent model.

Reading agent benchmarks without being misled

Vendor harnesses differ. Qwen evaluates TAU2-Bench with the airline-domain fixes from Anthropic’s Claude Opus 4.5 system card. Z.ai applied the same airline fixes and also added a prompt to the retail and telecom user simulations. Google reports Tau2 as an average over three without detailing its setup on the card. Each choice is reasonable, and together they mean a 79.1, a 79.5 and a 76.9 cannot be put in rank order.

Benchmark versions move. The τ-bench repository now ships τ³-bench, which adds a knowledge-retrieval banking domain, voice evaluation and more than 75 task fixes. Its July 2026 grading update states that results from versions before 1.0.1 are not comparable with later ones.

Single calls overstate agent skill. The Berkeley Function Calling Leaderboard, last updated 12 April 2026, scores single, parallel and multi-turn calling separately. The gap is large for most open models:

Model (BFCL, April 2026)Live single-call accuracyMulti-turn accuracy
Qwen3-32B (FC)82.01%47.87%
GLM-4.6 (FC, thinking)80.90%68.00%
Llama 3.3 70B Instruct (FC)76.61%21.50%
Gemma 3 27B (prompt)74.54%10.75%

The board predates Qwen3.5, Gemma 4 and gpt-oss, so use it for the pattern rather than a ranking of current models.

One harness for several models. NVIDIA’s Nemotron 3.5 Lightning card runs Qwen3.6-35B-A3B, Gemma 4 26B A4B, gpt-oss-20b and its own models through one harness. Qwen3.6-35B-A3B led most agentic rows there, including 48.74 on BrowseComp against 26.30 for Gemma 4 26B A4B, while Gemma 4 26B A4B led instruction following with 77.25 on IFBench.

Serving: parsers, strict mode and structured output

A model’s tool-call format only becomes a structured tool_calls response if the server parses it.

vLLM. Start the server with --enable-auto-tool-choice and the --tool-call-parser the model card names. vLLM’s tool calling documentation lists parsers including hermes, mistral, llama3_json, llama4_pythonic, openai for gpt-oss, glm45 and glm47, deepseek_v31, kimi_k2 and cohere_command4; the Qwen3.5 and Qwen3.6 cards specify qwen3_coder. Three settings decide how much you can trust the output:

  • tool_choice="required" and named function calls use structured outputs, so the response is guaranteed to parse against your schema.
  • With tool_choice="auto", vLLM constrains a tool’s arguments only when the tool sets strict: true, or when the operator raises --tool-strict-level. Its documentation notes that most OpenAI-compatible clients never set strict.
  • Otherwise, the docs warn, arguments may occasionally be malformed or violate the function’s schema.

llama.cpp. llama-server supports OpenAI-style function calling through Jinja chat templates, which current builds enable by default. Native handlers cover formats such as Llama 3.x, Hermes 2 and 3, Qwen 2.5, Mistral Nemo and GPT-OSS harmony; other templates fall back to a generic handler that, per the function-calling guide, may use more tokens. Parallel calls are off unless the request sets "parallel_tool_calls": true, and response_format accepts a JSON schema.

Ollama. The chat API takes a tools array and documents single, parallel and multi-turn agent loops. Structured outputs take a JSON schema in the format field, and Ollama advises repeating the schema in the prompt. Recent releases fixed Gemma 4 tool-response continuations and GLM tool calls dropped at the end of generation, so stay current.

For a wider comparison of these servers, see our vLLM, Ollama and llama.cpp guide.

Context budget for tool loops

Tool definitions, calls, results and reasoning all accumulate in the context, and every agent run active at once needs its own KV cache. Ollama’s documentation recommends at least 64,000 tokens for agents, and Qwen advises keeping at least 128K for Qwen3.6 to preserve thinking.

Worked example: parallel 64K agent runs on a 32 GB GPU. Using the method of our sizing calculator, 28.8 GB is usable. Subtract the weights file and divide by the cache one 64K-token run needs:

ModelWeightsCache per 64K runRuns that fit (16-bit)With an 8-bit cache
Qwen3.6-35B-A3B20.4 GB1.41 GB511
gpt-oss-20b12.1 GB1.61 GB1020
GLM-4.7-Flash18.2 GB3.55 GB25
Qwen3.8-27B19.0 GB4.45 GB24
Gemma 4 31B18.0 GB6.21 GB13

Qwen3.6-35B-A3B keeps a growing cache in only 10 of 40 layers, which is why it supports five runs where a dense 31B supports one. On 128 GB of unified memory, gpt-oss-120b fits 16 such runs, Qwen3.5-122B-A10B 18 and Mistral Small 4 22. The general sizing method is in our GPU requirements guide.

Two habits stretch the budget. Expose only the tools a task needs, because every schema is resent on every turn. And trim old tool output: Qwen’s own search agents drop earlier tool responses once their combined length passes a threshold, a strategy its cards call context folding.

Reasoning traces inside tool loops

Most current models think before they act, and each family wants its reasoning handled differently across turns:

  • Gemma 4. Google’s card says earlier thoughts must not be added to the history before the next user turn, except on tool-call turns, where thinking content should be preserved.
  • Qwen3.8. Preserved thinking is on by default, and reasoning_effort accepts low, medium and xhigh.
  • GLM-4.7-Flash. Z.ai turns on Preserved Thinking mode for its multi-turn agent benchmarks and serves it with the glm45 reasoning parser.
  • gpt-oss. OpenAI says the models work correctly only with its harmony response format, and reasoning effort can be set to low, medium or high.

Get this wrong and the model loses track of why it called a tool, or the context fills with stale reasoning. Use the chat template shipped with the checkpoint and the reasoning parser its card names, and test a full multi-turn loop rather than a single call. If you run a desktop agent harness rather than your own loop, our OpenClaw and Hermes Agent setup guide covers that configuration. For agents that mainly retrieve documents, see our local RAG model picks. If the agent has to run on a laptop or a 16 GB card, start from the small models ranked for tool calling.

How VDF AI fits

VDF AI Agents treats the model as one setting of an agent. Its product page lists connecting any LLM or small model per agent, including Ollama and other compatible endpoints, and assigning tools from an MCP-powered registry. Tool access is scoped per agent, high-risk actions can require human approval, and the audit log records every prompt, tool call and output.

The MCP gateway governs the tools themselves: administrators grant tools to roles in a registry, the MCP server layer runs inside your perimeter, including air-gapped estates, and each tool call leaves an audit record. That combination lets you swap the local model behind an agent, for example from Qwen3.6-35B-A3B to GLM-4.7-Flash, without changing which tools the agent may touch.

Sources

Model cards, leaderboards and documentation verified 6 October 2026.


Putting local models behind governed agents? See how VDF AI Agents scopes tools per agent, or book a demo.

Frequently asked questions

What is the best local LLM for agents?

On a 24 to 32 GB GPU, Qwen3.6-35B-A3B is the strongest general pick in October 2026: Apache 2.0, a 262K window, tool-calling support documented on its card, and only 3 billion active parameters, so its KV cache stays small across long tool loops. GLM-4.7-Flash, under MIT, is the main alternative. With 128 GB of memory, gpt-oss-120b, Mistral Small 4 and Qwen3.5-122B-A10B all fit with room for several parallel agent runs.

Which local LLM is best at tool calling?

Vendor scores point to the Qwen3.5 to 3.8 family, GLM-4.7-Flash and Gemma 4 31B: Qwen reports 79.1 on TAU2-Bench for Qwen3.5-9B, Z.ai reports 79.5 on τ²-Bench for GLM-4.7-Flash, and Google reports 76.9 percent on Tau2 for Gemma 4 31B. Each vendor ran its own harness with its own fixes, so the numbers cannot be ranked against each other. Run your candidates through your own tools before choosing.

What is the best small local model for tool calling?

Qwen3.5-9B is the strongest choice that fits an 8 GB graphics card, with Qwen reporting 66.1 on BFCL-V4 and 79.1 on TAU2-Bench. Qwen3.5-4B reports 79.9 on TAU2-Bench in the same table, at under 3 GB for a 4-bit file. Granite 4.2 8B and Ministral 3 8B both document native tool calling under Apache 2.0. Small models handle well-scoped tools well; long multi-step plans still favour larger models.

How much context does a local agent need?

More than a chat. Every turn appends the tool definitions, the model's call, the tool result and any reasoning, so one task can grow past 64,000 tokens. Ollama's documentation recommends at least 64,000 tokens for agents, and Qwen advises at least 128K for its Qwen3.6 models to preserve thinking. Size memory for the context of each run multiplied by the number of agent runs active at once, not for a single short chat.

Why does my local model output broken tool calls?

Usually the server is not parsing the model's native format. In vLLM you need the tool-call parser the model card names, such as qwen3_coder, glm47 or mistral, plus the auto tool-choice flag. vLLM's documentation warns that with automatic tool choice and no strict schemas, arguments may occasionally be malformed. Marking tools strict, using required or named tool choice, or raising the server's strictness level constrains the output to the declared schema.

Filed under
tool callingAI agentslocal LLMopen-weight modelsvLLMon-premises AI
Enterprise AI Agents

See enterprise AI agents in production

Watch how VDF AI runs governed, multi-agent workflows on your own infrastructure — then compare it against the platforms you are evaluating.

Or start free — no credit card →

Keep reading