Multimodal Tool

The Image Analysis Tool

Analyze an image to describe its content, read its text, and answer questions about it so an agent can reason over photos, screenshots, and diagrams — with vision that can run in your perimeter.

Explore VDF AI Agents
Any modalityText, speech, image, video
GovernedProcessed in your perimeter
AssignableTo any VDF AI agent
100%On-premise capable
The Locked-Media Problem

Most of your content isn’t text

Calls, recordings, scans, images, and foreign-language documents carry critical information that text-only agents simply can’t use. Turning that media into data usually means shipping sensitive content to a hosted API — exactly what regulated teams can’t do.

01

Media is opaque

Audio, video, and images are invisible to search and to agents.

02

Language barriers

Content in other languages stays out of reach.

03

Manual transcription

Transcribing and translating by hand is slow and costly.

04

Sensitive content

Calls and scans can’t be sent to a third-party service.

How the Tool Works

Image Analysis, without the risk

Capability

What it does

Understand what’s in an image.

it analyzes an image to describe content, read text, and answer questions about it.

Tool
Image Analysis

Assignable to any agent

VisionDescribeRead textOn-prem

How it works

Predictable, inspectable behavior

Designed to be reliable.

analysis can run on a vision model inside your perimeter, so an agent understands images — including sensitive ones — without sending them to a hosted API.

Governed
Policy + Audit

Every call logged

ScopedLoggedGovernedOn-prem

Governance

Private, governed, on-premise

Runs inside your perimeter.

Transcription, translation, and analysis can run on local models inside your perimeter with audit logging, so sensitive audio, video, and documents become usable data without ever leaving your environment.

100%
On-Prem

Per-tenant, logged

On-premRBACAudit logSovereign
Inputs

Parameters

The image_analyze tool accepts these inputs when an agent calls it. Required inputs are flagged.

Name Type Required Description
image_base64 string Optional Base64-encoded image.
image_url string Optional URL of the image (alternative to base64).
prompt string Optional Optional question to answer about the image.
In depth

How the Image Analysis tool works in practice

Most enterprise images cannot be sent to a hosted vision API: engineering drawings, patient scans, ID documents, internal dashboards, defect photos from the production line. The Image Analysis tool exists for exactly that material — it runs a vision model inside your perimeter, so an agent can describe a photo, read the text in a screenshot, or answer a targeted question about a diagram while the pixels never cross your network boundary.

In practice the prompt parameter does the heavy lifting. Ask nothing and you get a faithful description; ask “which value is highlighted in red in this table?” or “does this photo show corrosion on the flange?” and the answer comes back shaped for the workflow that asked. Paired with OCR Text Extraction for dense documents and Table Extraction for tabular scans, it gives the AI Document Analysis Agent working eyes.

Every call is logged like any other governed tool run — who invoked it, on which image, with which prompt and model — so visual analysis carries the same audit trail as text. Screenshots captured by Screenshot Capture can flow straight into it, which is how browsing agents verify what a page actually shows before acting on it.

Where it pays back

Where Image Analysis pays back

Screenshots

Understand a UI screenshot to act on it.

Documents

Read and interpret a scanned page or chart.

Inspection

Describe or classify a photo.

Accessibility

Generate alt text for images.

How VDF AI connects it

Assigned to agents, orchestrated as networks

On VDF AI, an industry’s use cases map to agents, and you assign tools like this one to those agents. Compose multiple agents into a governed, on-premise network.

  1. 1
    Industry Your sector Finance, healthcare, telecom, government, and more.
  2. 2
    Use Case A job to be done Concrete workflows the business needs solved.
  3. 3
    Agent A specialized worker Governed AI agents that execute the use case.
  4. 4
    Tool Image Analysis The capability you assign to an agent.
  5. 5
    Network Agents, orchestrated Many use cases and agents, working as one.
ROI Snapshot

What changes after you assign it

Unlocked
Media becomes searchable data
At-scale
Hours of media processed fast
Private
Sensitive media never leaves
100%
Runs on infrastructure you control
FAQ

Questions about the Image Analysis tool

What is the Image Analysis tool?

It analyzes an image to describe content, read text, and answer questions about it. Assigned to a VDF AI agent, it runs under role-based policy with full audit logging so the capability is safe to use in production.

Can it read text in the image?

Yes. It reads embedded text and can answer questions about the image content.

How is this different from OCR?

OCR extracts text; image analysis also describes, classifies, and reasons about the whole image.

Which image formats and sizes does it accept?

Standard raster formats — PNG, JPEG, WebP, and GIF frames — passed as base64 or a URL. Very large images are downscaled before inference, so pass the original resolution when small text must stay legible, or crop to the region that matters.

What does the analysis return?

A structured text answer shaped by your prompt: a scene description by default, or the specific fields you ask for — extracted text, object lists, classifications, or a yes/no judgment — ready for the next tool in the workflow to consume.

What inputs does the Image Analysis tool need?

It has no strictly required inputs, and optionally accepts image_base64, image_url, and prompt. Each parameter is validated when an agent calls the tool, and the full call is logged for audit.

Which tools pair well with Image Analysis?

Image Analysis is commonly assigned alongside Video Analysis, OCR Text Extraction, and Screenshot Capture. On VDF AI you compose several tools and agents into a single governed, on-premise network.

Does it run on-premise?

Yes. Like every VDF AI tool, it can run on-premise or in your sovereign cloud, scoped per user and audit-logged, so your data never leaves your perimeter.

How do agents use it?

You assign the tool to an agent under a role-based policy; the agent calls it as one step in a task, and several agents and tools can be orchestrated together as a governed VDF AI Network.

Put Image Analysis to work

See the Image Analysis tool assigned to an agent and orchestrated in a governed, on-premise network.

Try it free on VDF AI Deploy on your own infrastructure