The best local vision model for most teams in 2026 is a Qwen3.5 or Gemma 4 checkpoint sized to your GPU: both are Apache 2.0, read images natively and run on vLLM, llama.cpp and Ollama. Ministral 3 is Mistral's Apache 2.0 option, MiniCPM-V 4.6 runs on phones, and Llama 4 withholds its multimodal licence from EU-headquartered companies.
Through 2025, Qwen shipped vision as a separate VL line next to its text models. In 2026 vision is built into the mainstream open-weight families instead. Every Qwen3.5 size from 0.8B upward accepts images and video, Gemma 4 reads images at every size, and Mistral ships a vision encoder with each Ministral 3 model.
This guide compares the open-weight vision models that exist in October 2026, checked against their model cards and licences on Hugging Face, plus the support lists of vLLM, llama.cpp and Ollama. Benchmark numbers are the vendors’ own unless stated otherwise. Reading text from scanned pages is one use among several here; for document OCR pipelines, see our scanned-document RAG guide and on-premise document AI comparison.
Quick picks
| Job | Pick | Why |
|---|---|---|
| General image and video questions on one 16 GB GPU | Qwen3.5-9B | Image and video input, 262K context, published results on charts, documents and screens |
| Larger single-GPU deployments, 24–32 GB | Qwen3.5-27B or Gemma 4 31B at 4-bit | Apache 2.0, about 17–19 GB of weights at Q4_K_M |
| Laptops and edge boxes | Gemma 4 E4B or Qwen3.5-4B | Under 5 GB at 4-bit; Gemma’s E models also take audio |
| Phones | MiniCPM-V 4.6 | A 1.3B model with iOS, Android and HarmonyOS demos |
| Charts and tables into structured data | Granite Vision 4.1 4B | Chart-to-CSV, chart-to-code and table extraction modes |
| Screens and GUI agents | Qwen3.5-9B or GLM-4.6V-Flash | Screen benchmarks from Qwen; native multimodal function calling from Z.ai |
| Pointing and tracking in video | Molmo2-8B | Built for grounding in images and video, with published training data |
| A Mistral-based stack | Ministral 3 8B or 14B | 0.4B vision encoder, function calling, Apache 2.0 |
What local vision models handle well now
The model cards describe a wider range of work than captioning:
- Photos and inspection images. Gemma 4 lists object detection and pointing; Qwen3.5 adds spatial benchmarks such as RefCOCO and CountBench to its tables.
- Charts and dashboards. Gemma 4 names chart comprehension, and IBM’s Granite Vision 4.1 converts charts into CSV, summaries or Python code that redraws them.
- Screens and user interfaces. Qwen3-VL’s card describes operating PC and mobile interfaces, and Qwen3.5 publishes ScreenSpot Pro, OSWorld-Verified and AndroidWorld results.
- Video. Qwen3.5, Gemma 4, MiniCPM-V 4.6 and Molmo2 accept video, generally as sampled frames.
- Documents and handwriting. Gemma 4 lists document and PDF parsing, multilingual OCR and handwriting recognition among its capabilities.
- Code from images. Qwen3-VL’s card describes generating Draw.io diagrams and HTML, CSS and JavaScript from images or videos.
The shortlist: open-weight vision models in 2026
Checked against the Hugging Face model cards and licence files in October 2026.
| Model (vendor) | Sizes | Licence | Inputs | Context | Worth knowing |
|---|---|---|---|---|---|
| Qwen3.5 (Qwen) | 0.8B, 2B, 4B, 9B, 27B dense; 35B-A3B, 122B-A10B, 397B-A17B MoE | Apache 2.0 | Text, image, video | 262K | Vision in every size; Qwen3.6-27B and Qwen3.8-27B also take images |
| Qwen3-VL (Qwen) | 2B to 235B, dense and MoE | Apache 2.0 | Text, image, video | 256K, expandable to 1M | The 2025 dedicated VL line, with GUI-agent and visual-coding features |
| Gemma 4 (Google) | E2B, E4B, 12B, 26B A4B, 31B | Apache 2.0 | Text, image, video; audio on E2B, E4B and 12B | 128K on E2B and E4B; 256K on the rest | Visual token budgets from 70 to 1,120 per image; the 12B has no separate encoder |
| Ministral 3 (Mistral) | 3B, 8B, 14B | Apache 2.0 | Text, image | 256K | 0.4B vision encoder; fits 8, 12 and 24 GB of VRAM at FP8, per Mistral |
| Mistral Small 3.2 and Mistral Small 4 (Mistral) | 24B dense; 119B MoE with 6.5B active | Apache 2.0 | Text, image | 128K; 256K | Mistral publishes an NVFP4 build of Small 4 |
| Llama 4 Scout (Meta) | 109B MoE with 17B active | Llama 4 Community License, with an EU clause | Text, image | 10M | Meta says it fits one H100 with on-the-fly int4; tested with up to 5 images |
| InternVL3.5 (OpenGVLab) | 1B to 241B-A28B, including 8B, 14B, 38B and 30B-A3B | Apache 2.0 | Text, image | Not stated on the card | Qwen3 and gpt-oss language models with InternViT encoders |
| MiniCPM-V 4.6 (OpenBMB) | 1.3B | Apache 2.0 | Text, image, video | Not stated on the card | SigLIP2-400M encoder on Qwen3.5-0.8B |
| Molmo2 (Ai2) | 4B, 8B, O-7B | Apache 2.0 | Text, image, video | Not stated on the card | Grounding and pointing; trained on published datasets |
| GLM-4.6V-Flash and GLM-4.6V (Z.ai) | 9B and 106B | MIT | Text, image | 128K | Screenshots and pages can be passed straight to tools |
| Phi-4-reasoning-vision-15B (Microsoft) | 15B | MIT | Text, image | 16K | Reasoning backbone with a SigLIP-2 encoder; GUI grounding |
| Granite Vision 4.1 (IBM) | 4B | Apache 2.0 | Text, image | Not stated on the card | Chart, table and key-value extraction; works with Docling |
| LFM2.5-VL-3B (Liquid AI) | 3B | LFM Open License 1.0 | Text, image | 32K | Aimed at high-throughput, low-latency tasks on device |
Release dates span 2025 and 2026: Phi-4-reasoning-vision-15B’s card gives 4 March 2026 and Granite Vision 4.1’s gives 29 April 2026, while Qwen3-VL, InternVL3.5, Molmo2 and GLM-4.6V appeared on Hugging Face in the second half of 2025. Llama 4 Scout dates from April 2025, and Meta has released no newer Llama since. For text-first models under 15B that can also read images, see our small language model comparison.
Memory: weights, the vision encoder and image tokens
A vision model needs memory for three things: the language model’s weights, the vision encoder, and the image tokens that land in the KV cache. The weights below are Hugging Face parameter totals, which include the encoder, at three precisions. The arithmetic matches our VRAM calculator, with Q4_K_M at 4.89 bits per weight.
| Model | BF16 | FP8 | Q4_K_M | Vendor guidance |
|---|---|---|---|---|
| MiniCPM-V 4.6 (1.3B) | 2.6 GB | 1.3 GB | 0.8 GB | Runs on phones |
| Gemma 4 E4B (8.0B with embeddings) | 16.0 GB | 8.0 GB | 4.9 GB | Designed for laptops and mobile devices |
| Qwen3.5-9B | 19.3 GB | 9.7 GB | 5.9 GB | |
| Ministral 3 8B | 17.8 GB | 8.9 GB | 5.5 GB | Fits 12 GB of VRAM at FP8 |
| Molmo2-8B | 17.3 GB | 8.7 GB | 5.3 GB | |
| GLM-4.6V-Flash | 20.6 GB | 10.3 GB | 6.3 GB | Built for local, low-latency use |
| Phi-4-reasoning-vision-15B | 30.2 GB | 15.1 GB | 9.2 GB | |
| Mistral Small 3.2 24B | 48.0 GB | 24.0 GB | 14.7 GB | About 55 GB of GPU memory in BF16 |
| Qwen3.5-27B | 55.6 GB | 27.8 GB | 17.0 GB | |
| Gemma 4 31B | 62.5 GB | 31.3 GB | 19.1 GB | |
| Qwen3.5-35B-A3B | 71.9 GB | 36.0 GB | 22.0 GB | |
| Llama 4 Scout | 217.3 GB | 108.6 GB | 66.4 GB | One H100 with on-the-fly int4 |
Image tokens fill the KV cache. Gemma 4 makes the trade explicit with visual token budgets of 70, 140, 280, 560 or 1,120 tokens per image: Google advises lower budgets for classification, captioning and video frames, and higher ones for OCR, document parsing and small text. Ten images at 1,120 tokens add 11,200 tokens of context to one request. What that costs depends on the attention design, per each model’s config.json:
- Mistral Small 3.2 24B caches all 40 layers at 160 KiB per token, so those ten images take about 1.8 GB at 16 bits.
- Ministral 3 8B caches 34 layers at 136 KiB per token: about 1.6 GB.
- Gemma 4 12B keeps a full cache in only 8 of its 48 layers; the other 40 hold a 1,024-token window. The same request costs about 0.5 GB.
- Qwen3.5-9B caches only its 8 full-attention layers at 32 KiB per token: about 0.4 GB.
Multiply by the conversations you serve at once. When a deployment needs only text, vLLM’s --language-model-only flag skips the vision modules of Llama 4, Mistral 3 and Qwen3.5 models and frees that memory for the KV cache. Our GPU buyer’s guide maps these sizes to cards.
Serving support: vLLM, llama.cpp and Ollama
Checked against vLLM’s supported-model list, llama.cpp’s multimodal documentation, GGUF builds on Hugging Face from llama.cpp’s maintainers (ggml-org) and from model vendors, and Ollama’s library (October 2026).
| Model family | vLLM | llama.cpp | Ollama library |
|---|---|---|---|
| Qwen3.5 | Yes, dense and MoE | Yes; maintainer GGUF of 35B-A3B with a vision projector | Yes, qwen3.5, 0.8B to 122B |
| Qwen3-VL | Yes, dense and MoE | Yes; maintainer GGUFs for 2B and 30B-A3B | Yes, qwen3-vl |
| Gemma 4 | Yes, including the encoder-free 12B | Yes; listed in the multimodal docs | Yes, gemma4, E2B to 31B |
| Ministral 3 | Yes | Yes; maintainer GGUFs with a vision projector | Yes, ministral-3 |
| Mistral Small 3.2 | Yes | The docs list Mistral Small 3.1 | Yes, mistral-small3.2 |
| Llama 4 Scout | Yes | Yes; listed in the multimodal docs | Yes, llama4 |
| InternVL3.5 | Yes | The docs list InternVL 2.5 and 3 only | Not in the library |
| MiniCPM-V 4.6 | Yes | Yes; maintainer GGUF with a vision projector | Yes, minicpm-v4.6 |
| Molmo2 | Yes | Community GGUFs only; not in the docs | Not in the library |
| GLM-4.6V-Flash | Architecture listed | Yes; maintainer GGUF with a vision projector | Not in the library |
| Phi-4-reasoning-vision-15B | Yes | Community GGUFs only; not in the docs | Not in the library |
| Granite Vision 4.1 | Yes | IBM publishes a GGUF with a vision projector | Only the older granite3.2-vision |
| LFM2.5-VL-3B | Architecture listed | Liquid AI publishes a GGUF with a vision projector | Not in the library |
Three practical notes:
- llama.cpp loads vision as a second file. The model GGUF and a multimodal projector (
--mmproj) are separate, andllama-serveraccepts images through its OpenAI-compatible chat endpoint. Its docs warn that OCR-specialised models need their own prompt formats. - vLLM is the broadest server. It lists every family in the table, which makes it the default for shared GPU servers and for models Ollama does not carry.
- Ollama is the quickest start for the families it lists. Check the vision tag on the library page before you pull a model.
Matching a model to images, charts, screens and video
Vendor-reported figures, useful for shortlisting rather than ranking across vendors:
- Photos and general questions. Qwen reports 80.3 on RealWorldQA and 79.7 on MMStar for Qwen3.5-9B. Gemma 4 adds pointing and object detection, and Molmo2 is built around grounding answers in specific image regions.
- Charts and dashboards. Granite Vision 4.1 has dedicated modes that turn a chart into CSV, a summary or Python code, and extracts tables to JSON, HTML or OTSL. For question answering over charts, Qwen reports 73.0 on CharXiv reasoning questions for Qwen3.5-9B.
- Screens and GUI agents. Qwen3.5-9B scores 65.2 on ScreenSpot Pro, 41.8 on OSWorld-Verified and 57.8 on AndroidWorld, per Qwen. GLM-4.6V-Flash accepts screenshots and rendered pages as direct tool inputs through native multimodal function calling. Phi-4-reasoning-vision-15B is tuned for GUI grounding, with ScreenSpot-V2 among its reported evaluations.
- Video. Qwen reports 84.5 on VideoMME with subtitles for Qwen3.5-9B. Molmo2’s training sets include video pointing and tracking data, and MiniCPM-V 4.6 handles video on a phone-sized model.
- Pages and forms. Most models above read printed text, and Qwen reports 89.2 on OCRBench for Qwen3.5-9B. Production pipelines for scanned documents also need layout analysis, chunking and retrieval, which the document guides linked above cover.
Benchmarks rarely match your images. Build a test set from your own screenshots, charts and photos, at the resolutions you will actually send, and compare two candidates before standardising.
Licences: mostly Apache and MIT, with two exceptions
Qwen3.5, Qwen3-VL, Gemma 4, Ministral 3, Mistral Small 3.2 and 4, InternVL3.5, MiniCPM-V 4.6, Molmo2 and Granite Vision 4.1 use Apache 2.0. GLM-4.6V and Phi-4-reasoning-vision-15B use MIT. Two licences need care:
- Llama 4. Meta’s acceptable use policy withholds the Section 1(a) licence rights for Llama 4’s multimodal models from individuals domiciled in the EU and from companies whose principal place of business is there. End users of a product that incorporates the models are not affected. Scout and Maverick both accept images, so an EU-headquartered company cannot self-host either under the standard grant. Llama 3.2’s policy applies the same clause to its multimodal models. Our Llama licence guide covers the detail.
- LFM2.5-VL. The LFM Open License 1.0 grants commercial rights only to organisations below $10 million in annual revenue, with an exception for qualifying non-profits doing research.
Molmo2 adds a provenance point in its favour: Ai2 says it was trained on publicly available third-party datasets, which it lists in its technical report and on Hugging Face.
How VDF AI fits
VDF AI exposes vision to agents through two tools. Image Analysis describes an image, reads its text and answers questions about it, for screenshots, scanned pages, charts and inspection photos. Video Analysis samples frames to describe scenes, objects and actions. Both can run on a vision model inside your perimeter, so sensitive images and footage need not go to a hosted API.
VDF AI Router registers Ollama and custom on-premises deployments alongside any cloud models your policy permits, and its air-gap mode keeps every request on local models. A Qwen3.5 or Gemma 4 endpoint on your own GPU can therefore sit behind the same governed routing as your text models.
Sources
Verified 6 October 2026.
- Qwen3.5-9B model card and Qwen3.5-0.8B model card
- Qwen3.5-27B and Qwen3.5-35B-A3B model cards
- Qwen3-VL-8B-Instruct model card
- Gemma 4 E4B model card, Gemma 4 12B and Gemma 4 31B
- Ministral 3 8B model card
- Mistral Small 3.2 model card, Mistral Small 4 model card and its NVFP4 build
- Llama 4 model card and acceptable use policy
- Llama 3.2 acceptable use policy
- InternVL3.5-8B model card
- MiniCPM-V 4.6 model card
- Molmo2-8B model card
- GLM-4.6V-Flash model card
- Phi-4-reasoning-vision-15B model card
- Granite Vision 4.1 4B model card
- LFM2.5-VL-3B model card and LFM Open License 1.0
- vLLM supported models
- llama.cpp multimodal documentation
- ggml-org multimodal GGUF builds on Hugging Face, plus vendor builds from IBM and Liquid AI
- Ollama vision models
- Mistral Small 3.2 config.json and Gemma 4 12B config.json
Keeping images and screenshots inside your network? See the Image Analysis tool or book a demo.