AI Infrastructure

Best Local Vision Models (2026): Open-Weight VLMs for Images, Charts and Screens

Open-weight vision language models you can run on your own GPUs, checked against their model cards in October 2026: sizes, licences including Llama 4's EU clause, memory for weights and image tokens, serving support in vLLM, llama.cpp and Ollama, and which model suits photos, charts, screens or video.

The best local vision model for most teams in 2026 is a Qwen3.5 or Gemma 4 checkpoint sized to your GPU: both are Apache 2.0, read images natively and run on vLLM, llama.cpp and Ollama. Ministral 3 is Mistral's Apache 2.0 option, MiniCPM-V 4.6 runs on phones, and Llama 4 withholds its multimodal licence from EU-headquartered companies.

Through 2025, Qwen shipped vision as a separate VL line next to its text models. In 2026 vision is built into the mainstream open-weight families instead. Every Qwen3.5 size from 0.8B upward accepts images and video, Gemma 4 reads images at every size, and Mistral ships a vision encoder with each Ministral 3 model.

This guide compares the open-weight vision models that exist in October 2026, checked against their model cards and licences on Hugging Face, plus the support lists of vLLM, llama.cpp and Ollama. Benchmark numbers are the vendors’ own unless stated otherwise. Reading text from scanned pages is one use among several here; for document OCR pipelines, see our scanned-document RAG guide and on-premise document AI comparison.

Quick picks

JobPickWhy
General image and video questions on one 16 GB GPUQwen3.5-9BImage and video input, 262K context, published results on charts, documents and screens
Larger single-GPU deployments, 24–32 GBQwen3.5-27B or Gemma 4 31B at 4-bitApache 2.0, about 17–19 GB of weights at Q4_K_M
Laptops and edge boxesGemma 4 E4B or Qwen3.5-4BUnder 5 GB at 4-bit; Gemma’s E models also take audio
PhonesMiniCPM-V 4.6A 1.3B model with iOS, Android and HarmonyOS demos
Charts and tables into structured dataGranite Vision 4.1 4BChart-to-CSV, chart-to-code and table extraction modes
Screens and GUI agentsQwen3.5-9B or GLM-4.6V-FlashScreen benchmarks from Qwen; native multimodal function calling from Z.ai
Pointing and tracking in videoMolmo2-8BBuilt for grounding in images and video, with published training data
A Mistral-based stackMinistral 3 8B or 14B0.4B vision encoder, function calling, Apache 2.0

What local vision models handle well now

The model cards describe a wider range of work than captioning:

  • Photos and inspection images. Gemma 4 lists object detection and pointing; Qwen3.5 adds spatial benchmarks such as RefCOCO and CountBench to its tables.
  • Charts and dashboards. Gemma 4 names chart comprehension, and IBM’s Granite Vision 4.1 converts charts into CSV, summaries or Python code that redraws them.
  • Screens and user interfaces. Qwen3-VL’s card describes operating PC and mobile interfaces, and Qwen3.5 publishes ScreenSpot Pro, OSWorld-Verified and AndroidWorld results.
  • Video. Qwen3.5, Gemma 4, MiniCPM-V 4.6 and Molmo2 accept video, generally as sampled frames.
  • Documents and handwriting. Gemma 4 lists document and PDF parsing, multilingual OCR and handwriting recognition among its capabilities.
  • Code from images. Qwen3-VL’s card describes generating Draw.io diagrams and HTML, CSS and JavaScript from images or videos.

The shortlist: open-weight vision models in 2026

Checked against the Hugging Face model cards and licence files in October 2026.

Model (vendor)SizesLicenceInputsContextWorth knowing
Qwen3.5 (Qwen)0.8B, 2B, 4B, 9B, 27B dense; 35B-A3B, 122B-A10B, 397B-A17B MoEApache 2.0Text, image, video262KVision in every size; Qwen3.6-27B and Qwen3.8-27B also take images
Qwen3-VL (Qwen)2B to 235B, dense and MoEApache 2.0Text, image, video256K, expandable to 1MThe 2025 dedicated VL line, with GUI-agent and visual-coding features
Gemma 4 (Google)E2B, E4B, 12B, 26B A4B, 31BApache 2.0Text, image, video; audio on E2B, E4B and 12B128K on E2B and E4B; 256K on the restVisual token budgets from 70 to 1,120 per image; the 12B has no separate encoder
Ministral 3 (Mistral)3B, 8B, 14BApache 2.0Text, image256K0.4B vision encoder; fits 8, 12 and 24 GB of VRAM at FP8, per Mistral
Mistral Small 3.2 and Mistral Small 4 (Mistral)24B dense; 119B MoE with 6.5B activeApache 2.0Text, image128K; 256KMistral publishes an NVFP4 build of Small 4
Llama 4 Scout (Meta)109B MoE with 17B activeLlama 4 Community License, with an EU clauseText, image10MMeta says it fits one H100 with on-the-fly int4; tested with up to 5 images
InternVL3.5 (OpenGVLab)1B to 241B-A28B, including 8B, 14B, 38B and 30B-A3BApache 2.0Text, imageNot stated on the cardQwen3 and gpt-oss language models with InternViT encoders
MiniCPM-V 4.6 (OpenBMB)1.3BApache 2.0Text, image, videoNot stated on the cardSigLIP2-400M encoder on Qwen3.5-0.8B
Molmo2 (Ai2)4B, 8B, O-7BApache 2.0Text, image, videoNot stated on the cardGrounding and pointing; trained on published datasets
GLM-4.6V-Flash and GLM-4.6V (Z.ai)9B and 106BMITText, image128KScreenshots and pages can be passed straight to tools
Phi-4-reasoning-vision-15B (Microsoft)15BMITText, image16KReasoning backbone with a SigLIP-2 encoder; GUI grounding
Granite Vision 4.1 (IBM)4BApache 2.0Text, imageNot stated on the cardChart, table and key-value extraction; works with Docling
LFM2.5-VL-3B (Liquid AI)3BLFM Open License 1.0Text, image32KAimed at high-throughput, low-latency tasks on device

Release dates span 2025 and 2026: Phi-4-reasoning-vision-15B’s card gives 4 March 2026 and Granite Vision 4.1’s gives 29 April 2026, while Qwen3-VL, InternVL3.5, Molmo2 and GLM-4.6V appeared on Hugging Face in the second half of 2025. Llama 4 Scout dates from April 2025, and Meta has released no newer Llama since. For text-first models under 15B that can also read images, see our small language model comparison.

Memory: weights, the vision encoder and image tokens

A vision model needs memory for three things: the language model’s weights, the vision encoder, and the image tokens that land in the KV cache. The weights below are Hugging Face parameter totals, which include the encoder, at three precisions. The arithmetic matches our VRAM calculator, with Q4_K_M at 4.89 bits per weight.

ModelBF16FP8Q4_K_MVendor guidance
MiniCPM-V 4.6 (1.3B)2.6 GB1.3 GB0.8 GBRuns on phones
Gemma 4 E4B (8.0B with embeddings)16.0 GB8.0 GB4.9 GBDesigned for laptops and mobile devices
Qwen3.5-9B19.3 GB9.7 GB5.9 GB
Ministral 3 8B17.8 GB8.9 GB5.5 GBFits 12 GB of VRAM at FP8
Molmo2-8B17.3 GB8.7 GB5.3 GB
GLM-4.6V-Flash20.6 GB10.3 GB6.3 GBBuilt for local, low-latency use
Phi-4-reasoning-vision-15B30.2 GB15.1 GB9.2 GB
Mistral Small 3.2 24B48.0 GB24.0 GB14.7 GBAbout 55 GB of GPU memory in BF16
Qwen3.5-27B55.6 GB27.8 GB17.0 GB
Gemma 4 31B62.5 GB31.3 GB19.1 GB
Qwen3.5-35B-A3B71.9 GB36.0 GB22.0 GB
Llama 4 Scout217.3 GB108.6 GB66.4 GBOne H100 with on-the-fly int4

Image tokens fill the KV cache. Gemma 4 makes the trade explicit with visual token budgets of 70, 140, 280, 560 or 1,120 tokens per image: Google advises lower budgets for classification, captioning and video frames, and higher ones for OCR, document parsing and small text. Ten images at 1,120 tokens add 11,200 tokens of context to one request. What that costs depends on the attention design, per each model’s config.json:

  • Mistral Small 3.2 24B caches all 40 layers at 160 KiB per token, so those ten images take about 1.8 GB at 16 bits.
  • Ministral 3 8B caches 34 layers at 136 KiB per token: about 1.6 GB.
  • Gemma 4 12B keeps a full cache in only 8 of its 48 layers; the other 40 hold a 1,024-token window. The same request costs about 0.5 GB.
  • Qwen3.5-9B caches only its 8 full-attention layers at 32 KiB per token: about 0.4 GB.

Multiply by the conversations you serve at once. When a deployment needs only text, vLLM’s --language-model-only flag skips the vision modules of Llama 4, Mistral 3 and Qwen3.5 models and frees that memory for the KV cache. Our GPU buyer’s guide maps these sizes to cards.

Serving support: vLLM, llama.cpp and Ollama

Checked against vLLM’s supported-model list, llama.cpp’s multimodal documentation, GGUF builds on Hugging Face from llama.cpp’s maintainers (ggml-org) and from model vendors, and Ollama’s library (October 2026).

Model familyvLLMllama.cppOllama library
Qwen3.5Yes, dense and MoEYes; maintainer GGUF of 35B-A3B with a vision projectorYes, qwen3.5, 0.8B to 122B
Qwen3-VLYes, dense and MoEYes; maintainer GGUFs for 2B and 30B-A3BYes, qwen3-vl
Gemma 4Yes, including the encoder-free 12BYes; listed in the multimodal docsYes, gemma4, E2B to 31B
Ministral 3YesYes; maintainer GGUFs with a vision projectorYes, ministral-3
Mistral Small 3.2YesThe docs list Mistral Small 3.1Yes, mistral-small3.2
Llama 4 ScoutYesYes; listed in the multimodal docsYes, llama4
InternVL3.5YesThe docs list InternVL 2.5 and 3 onlyNot in the library
MiniCPM-V 4.6YesYes; maintainer GGUF with a vision projectorYes, minicpm-v4.6
Molmo2YesCommunity GGUFs only; not in the docsNot in the library
GLM-4.6V-FlashArchitecture listedYes; maintainer GGUF with a vision projectorNot in the library
Phi-4-reasoning-vision-15BYesCommunity GGUFs only; not in the docsNot in the library
Granite Vision 4.1YesIBM publishes a GGUF with a vision projectorOnly the older granite3.2-vision
LFM2.5-VL-3BArchitecture listedLiquid AI publishes a GGUF with a vision projectorNot in the library

Three practical notes:

  • llama.cpp loads vision as a second file. The model GGUF and a multimodal projector (--mmproj) are separate, and llama-server accepts images through its OpenAI-compatible chat endpoint. Its docs warn that OCR-specialised models need their own prompt formats.
  • vLLM is the broadest server. It lists every family in the table, which makes it the default for shared GPU servers and for models Ollama does not carry.
  • Ollama is the quickest start for the families it lists. Check the vision tag on the library page before you pull a model.

Matching a model to images, charts, screens and video

Vendor-reported figures, useful for shortlisting rather than ranking across vendors:

  • Photos and general questions. Qwen reports 80.3 on RealWorldQA and 79.7 on MMStar for Qwen3.5-9B. Gemma 4 adds pointing and object detection, and Molmo2 is built around grounding answers in specific image regions.
  • Charts and dashboards. Granite Vision 4.1 has dedicated modes that turn a chart into CSV, a summary or Python code, and extracts tables to JSON, HTML or OTSL. For question answering over charts, Qwen reports 73.0 on CharXiv reasoning questions for Qwen3.5-9B.
  • Screens and GUI agents. Qwen3.5-9B scores 65.2 on ScreenSpot Pro, 41.8 on OSWorld-Verified and 57.8 on AndroidWorld, per Qwen. GLM-4.6V-Flash accepts screenshots and rendered pages as direct tool inputs through native multimodal function calling. Phi-4-reasoning-vision-15B is tuned for GUI grounding, with ScreenSpot-V2 among its reported evaluations.
  • Video. Qwen reports 84.5 on VideoMME with subtitles for Qwen3.5-9B. Molmo2’s training sets include video pointing and tracking data, and MiniCPM-V 4.6 handles video on a phone-sized model.
  • Pages and forms. Most models above read printed text, and Qwen reports 89.2 on OCRBench for Qwen3.5-9B. Production pipelines for scanned documents also need layout analysis, chunking and retrieval, which the document guides linked above cover.

Benchmarks rarely match your images. Build a test set from your own screenshots, charts and photos, at the resolutions you will actually send, and compare two candidates before standardising.

Licences: mostly Apache and MIT, with two exceptions

Qwen3.5, Qwen3-VL, Gemma 4, Ministral 3, Mistral Small 3.2 and 4, InternVL3.5, MiniCPM-V 4.6, Molmo2 and Granite Vision 4.1 use Apache 2.0. GLM-4.6V and Phi-4-reasoning-vision-15B use MIT. Two licences need care:

  • Llama 4. Meta’s acceptable use policy withholds the Section 1(a) licence rights for Llama 4’s multimodal models from individuals domiciled in the EU and from companies whose principal place of business is there. End users of a product that incorporates the models are not affected. Scout and Maverick both accept images, so an EU-headquartered company cannot self-host either under the standard grant. Llama 3.2’s policy applies the same clause to its multimodal models. Our Llama licence guide covers the detail.
  • LFM2.5-VL. The LFM Open License 1.0 grants commercial rights only to organisations below $10 million in annual revenue, with an exception for qualifying non-profits doing research.

Molmo2 adds a provenance point in its favour: Ai2 says it was trained on publicly available third-party datasets, which it lists in its technical report and on Hugging Face.

How VDF AI fits

VDF AI exposes vision to agents through two tools. Image Analysis describes an image, reads its text and answers questions about it, for screenshots, scanned pages, charts and inspection photos. Video Analysis samples frames to describe scenes, objects and actions. Both can run on a vision model inside your perimeter, so sensitive images and footage need not go to a hosted API.

VDF AI Router registers Ollama and custom on-premises deployments alongside any cloud models your policy permits, and its air-gap mode keeps every request on local models. A Qwen3.5 or Gemma 4 endpoint on your own GPU can therefore sit behind the same governed routing as your text models.

Sources

Verified 6 October 2026.


Keeping images and screenshots inside your network? See the Image Analysis tool or book a demo.

Frequently asked questions

What is the best local vision model?

It depends on your GPU and your images. On a single 16 GB card, Qwen3.5-9B is the strongest all-rounder among the models checked here, with image and video input and published results on charts, documents and screen tasks. With 24 to 32 GB, Qwen3.5-27B or Gemma 4 31B at 4-bit go further. On laptops, Gemma 4 E4B or Qwen3.5-4B fit easily, and MiniCPM-V 4.6 runs on phones. For turning charts and tables into data, IBM's Granite Vision 4.1 is built for exactly that job.

Can Ollama run vision models locally?

Yes. Ollama's library marks qwen3.5, gemma4, ministral-3, qwen3-vl, llama4, mistral-small3.2 and minicpm-v4.6 as vision models, among others, so a pull and a prompt with an image path is enough to start. Several strong research models are not in the library, including InternVL3.5, Molmo2, Phi-4-reasoning-vision and LFM2.5-VL. For those, use vLLM, which lists all four architectures, or Hugging Face Transformers. llama.cpp also serves many vision models when you load the separate projector file.

What GPU do I need for a local vision model?

Plan for the weights, the vision encoder and the image tokens. Qwen3.5-9B takes about 5.9 GB at the Q4_K_M GGUF setting, and Mistral says Ministral 3 8B fits in 12 GB of VRAM at FP8. Each image becomes hundreds to over a thousand tokens of context: Gemma 4 lets you choose budgets from 70 to 1,120 tokens per image. Ten detailed images on Mistral Small 3.2 add about 1.8 GB of KV cache. A 16 GB card covers 8 to 9B vision models; 24 to 32 GB opens up the 27 to 31B tier.

Can EU companies use Llama 4 for vision?

Not under the standard licence, read literally. Meta's Llama 4 acceptable use policy says the rights in Section 1(a) of the licence are not granted for its multimodal models to individuals domiciled in the EU or to companies with their principal place of business there. Scout and Maverick both accept images, so both are covered. End users of a product built on the models are exempt. EU-headquartered teams usually choose Apache 2.0 alternatives such as Qwen3.5, Gemma 4 or Ministral 3, and should ask counsel about their group structure.

Which local vision model is best for screenshots and GUI agents?

Start with Qwen3.5-9B. Qwen reports 65.2 on ScreenSpot Pro, 41.8 on OSWorld-Verified and 57.8 on AndroidWorld for it, and its earlier Qwen3-VL line was built to operate PC and mobile interfaces. GLM-4.6V-Flash from Z.ai accepts screenshots and document pages directly as tool inputs through native multimodal function calling, under an MIT licence. Microsoft's Phi-4-reasoning-vision-15B is tuned for GUI grounding. Agent results depend heavily on the harness, so test on your own applications and screen resolutions.

Filed under
vision language modelsopen-weight modelslocal LLMon-premises AImodel servingvLLM
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading