The Image Analysis Tool
Analyze an image to describe its content, read its text, and answer questions about it so an agent can reason over photos, screenshots, and diagrams — with vision that can run in your perimeter.
Most of your content isn’t text
Calls, recordings, scans, images, and foreign-language documents carry critical information that text-only agents simply can’t use. Turning that media into data usually means shipping sensitive content to a hosted API — exactly what regulated teams can’t do.
Media is opaque
Audio, video, and images are invisible to search and to agents.
Language barriers
Content in other languages stays out of reach.
Manual transcription
Transcribing and translating by hand is slow and costly.
Sensitive content
Calls and scans can’t be sent to a third-party service.
Image Analysis, without the risk
Capability
What it does
Understand what’s in an image.
it analyzes an image to describe content, read text, and answer questions about it.
Assignable to any agent
How it works
Predictable, inspectable behavior
Designed to be reliable.
analysis can run on a vision model inside your perimeter, so an agent understands images — including sensitive ones — without sending them to a hosted API.
Every call logged
Governance
Private, governed, on-premise
Runs inside your perimeter.
Transcription, translation, and analysis can run on local models inside your perimeter with audit logging, so sensitive audio, video, and documents become usable data without ever leaving your environment.
Per-tenant, logged
Parameters
The image_analyze tool accepts these inputs when an agent calls it. Required inputs are flagged.
How the Image Analysis tool works in practice
Most enterprise images cannot be sent to a hosted vision API: engineering drawings, patient scans, ID documents, internal dashboards, defect photos from the production line. The Image Analysis tool exists for exactly that material — it runs a vision model inside your perimeter, so an agent can describe a photo, read the text in a screenshot, or answer a targeted question about a diagram while the pixels never cross your network boundary.
In practice the prompt parameter does the heavy lifting. Ask nothing and you get a faithful description; ask “which value is highlighted in red in this table?” or “does this photo show corrosion on the flange?” and the answer comes back shaped for the workflow that asked. Paired with OCR Text Extraction for dense documents and Table Extraction for tabular scans, it gives the AI Document Analysis Agent working eyes.
Every call is logged like any other governed tool run — who invoked it, on which image, with which prompt and model — so visual analysis carries the same audit trail as text. Screenshots captured by Screenshot Capture can flow straight into it, which is how browsing agents verify what a page actually shows before acting on it.
Where Image Analysis pays back
Screenshots
Understand a UI screenshot to act on it.
Documents
Read and interpret a scanned page or chart.
Inspection
Describe or classify a photo.
Accessibility
Generate alt text for images.
Assigned to agents, orchestrated as networks
On VDF AI, an industry’s use cases map to agents, and you assign tools like this one to those agents. Compose multiple agents into a governed, on-premise network.
- 1Industry Your sector Finance, healthcare, telecom, government, and more.
- 2Use Case A job to be done Concrete workflows the business needs solved.
- 3Agent A specialized worker Governed AI agents that execute the use case.
- 4Tool Image Analysis The capability you assign to an agent.
- 5Network Agents, orchestrated Many use cases and agents, working as one.
What changes after you assign it
Questions about the Image Analysis tool
What is the Image Analysis tool?
It analyzes an image to describe content, read text, and answer questions about it. Assigned to a VDF AI agent, it runs under role-based policy with full audit logging so the capability is safe to use in production.
Can it read text in the image?
Yes. It reads embedded text and can answer questions about the image content.
How is this different from OCR?
OCR extracts text; image analysis also describes, classifies, and reasons about the whole image.
Which image formats and sizes does it accept?
Standard raster formats — PNG, JPEG, WebP, and GIF frames — passed as base64 or a URL. Very large images are downscaled before inference, so pass the original resolution when small text must stay legible, or crop to the region that matters.
What does the analysis return?
A structured text answer shaped by your prompt: a scene description by default, or the specific fields you ask for — extracted text, object lists, classifications, or a yes/no judgment — ready for the next tool in the workflow to consume.
What inputs does the Image Analysis tool need?
It has no strictly required inputs, and optionally accepts image_base64, image_url, and prompt. Each parameter is validated when an agent calls the tool, and the full call is logged for audit.
Which tools pair well with Image Analysis?
Image Analysis is commonly assigned alongside Video Analysis, OCR Text Extraction, and Screenshot Capture. On VDF AI you compose several tools and agents into a single governed, on-premise network.
Does it run on-premise?
Yes. Like every VDF AI tool, it can run on-premise or in your sovereign cloud, scoped per user and audit-logged, so your data never leaves your perimeter.
How do agents use it?
You assign the tool to an agent under a role-based policy; the agent calls it as one step in a task, and several agents and tools can be orchestrated together as a governed VDF AI Network.
Assign Image Analysis to these agents
These VDF AI agents can be assigned this tool. Open an agent to see the full toolkit it can run.
Tools that work well alongside this one
Where this tool delivers value
Put Image Analysis to work
See the Image Analysis tool assigned to an agent and orchestrated in a governed, on-premise network.