The Screenshot Capture Tool
Take a screenshot of the full page or a specific element so an agent can verify what a page looks like, keep visual evidence, or feed the image to vision analysis.
Half the work lives behind a browser
Critical data and actions sit on web pages and third-party systems with no clean API. Without the ability to navigate, click, extract, and call endpoints, an agent is cut off from a huge share of the real work — and doing it by hand doesn’t scale.
No API, no access
Many systems only expose a web UI an agent can’t reach.
Manual gathering
People copy-paste from pages and PDFs because nothing else can.
Brittle scripts
Hand-written scrapers break the moment a page changes.
Ungoverned egress
Uncontrolled outbound web calls are a security and compliance risk.
Screenshot Capture, without the risk
Capability
What it does
Capture a visual image of the page.
it captures a screenshot of the full page or a specified element.
Assignable to any agent
How it works
Predictable, inspectable behavior
Designed to be reliable.
captures run in the governed browser and are stored in your environment, so an agent produces visual evidence and inputs for analysis without anything leaving your perimeter.
Every call logged
Governance
Private, governed, on-premise
Runs inside your perimeter.
Browsing and web calls run through a governed, on-premise gateway with allow-listing and full request logging, so agents reach only the sites and endpoints you permit — and every action is auditable.
Per-tenant, logged
Parameters
The screenshot_capture tool accepts these inputs when an agent calls it. Required inputs are flagged.
default: false Optional Capture the entire scrollable page.
How the Screenshot Capture tool works in practice
An agent that operates a browser needs proof of what it saw. The Screenshot Capture tool produces that proof: a pixel-accurate record of the page — full page, viewport, or a single element — taken inside the governed browser session and stored in your environment. When a workflow later gets audited, “what did the portal actually display before the agent clicked submit?” has a visual answer.
Three patterns dominate. Visual QA: capture after a deploy or content change and let Image Analysis judge whether the page rendered correctly. Evidence: staple the screenshot to the run record of an irreversible action, next to the structured page state from Browser Snapshot. Reporting: drop captures into the documents that Report Compile assembles, so stakeholders see the page, not a paraphrase of it.
Captures never leave your perimeter — relevant when the pages being photographed are logged-in internal systems — and each one is tied to the agent, session, and URL that produced it, which is precisely what separates governed browsing from a scraper with a screenshot habit.
Where Screenshot Capture pays back
Visual QA
Confirm a page rendered as expected.
Evidence
Keep a visual record of a state or result.
Vision input
Feed a screenshot to image analysis.
Reporting
Include page images in an agent’s report.
Assigned to agents, orchestrated as networks
On VDF AI, an industry’s use cases map to agents, and you assign tools like this one to those agents. Compose multiple agents into a governed, on-premise network.
- 1Industry Your sector Finance, healthcare, telecom, government, and more.
- 2Use Case A job to be done Concrete workflows the business needs solved.
- 3Agent A specialized worker Governed AI agents that execute the use case.
- 4Tool Screenshot Capture The capability you assign to an agent.
- 5Network Agents, orchestrated Many use cases and agents, working as one.
What changes after you assign it
Questions about the Screenshot Capture tool
What is the Screenshot Capture tool?
It captures a screenshot of the full page or a specified element. Assigned to a VDF AI agent, it runs under role-based policy with full audit logging so the capability is safe to use in production.
Can it capture the whole page?
Yes. Set full_page to true, or pass a selector to capture a single element.
Where is the image stored?
In your own environment; nothing is sent to a third party.
Full page, viewport, or a single element — what exactly gets captured?
Your choice per call: the visible viewport by default, the entire scrollable page with full_page: true, or one element when you pass its selector — useful for capturing just the error banner, the chart, or the form that matters.
What format is the output and where does it go?
A PNG image stored in your own environment and attached to the agent run that produced it. From there it can be embedded in a report, kept as evidence, or handed to Image Analysis for a model to interpret what the page showed.
What inputs does the Screenshot Capture tool need?
It has no strictly required inputs, and optionally accepts selector and full_page. Each parameter is validated when an agent calls the tool, and the full call is logged for audit.
Which tools pair well with Screenshot Capture?
Screenshot Capture is commonly assigned alongside Browser Snapshot, Image Analysis, and Browser Navigate. On VDF AI you compose several tools and agents into a single governed, on-premise network.
Does it run on-premise?
Yes. Like every VDF AI tool, it can run on-premise or in your sovereign cloud, scoped per user and audit-logged, so your data never leaves your perimeter.
How do agents use it?
You assign the tool to an agent under a role-based policy; the agent calls it as one step in a task, and several agents and tools can be orchestrated together as a governed VDF AI Network.
Assign Screenshot Capture to these agents
These VDF AI agents can be assigned this tool. Open an agent to see the full toolkit it can run.
Tools that work well alongside this one
Where this tool delivers value
Put Screenshot Capture to work
See the Screenshot Capture tool assigned to an agent and orchestrated in a governed, on-premise network.