AI Internal Audit Agent Risk & Assurance Agents Tier 2 On-premise Updated September 2026
AI Internal Audit Agent

AI Agent for Internal Audit Fieldwork

Sampling was never a methodology, it was a concession to how long reading takes. This agent tests the full population against your audit programme, records the document behind each exception, and hands the auditor a drafted workpaper instead of a spreadsheet to start from.

Population Tested in full rather than by sample
Sourced Every exception cites the document behind it
Drafted Workpapers prepared for auditor review
Auditor Opinions and ratings stay with the auditor
Examines
Transaction records Approval evidence Policies and procedures Control descriptions Prior findings Audit programmes

What is an AI internal audit agent?

An AI internal audit agent is a governed software worker that carries out internal audit fieldwork. It tests a full population against an organisation’s own audit programme instead of a sample, retains the record establishing each exception, identifies findings that recur across engagements, and drafts structured workpapers for an auditor to review.

What it does

Tests the full population, not a sample Applies your own audit programme criteria Keeps the record behind each exception Finds control failures recurring across audits Drafts workpapers in your own format

What it is not

Not the audit opinion or rating Not a change to any tested record Not a substitute for auditor judgement
The Coverage Problem

A five percent sample and a finding nobody can reproduce

Internal audit is asked for assurance over a population it can only read a fraction of, so a clean sample becomes an opinion about the whole. Meanwhile the fieldwork that did happen is recorded as a conclusion rather than as evidence, and when the finding is challenged six months later, reconstructing why it was reached is its own project.

Sampling misses the pattern

Twenty-five items out of forty thousand cannot detect behaviour that concentrates in a small, deliberate corner of the population.

Evidence is summarised away

A workpaper records that approval was obtained without keeping the record that establishes it, so the test is not repeatable.

Fieldwork consumes the audit

Most of an engagement is pulling documents and reconciling lists, leaving the least time for the judgement the auditor is there to apply.

Findings recur without being seen

The same control fails in three consecutive audits of different business units and nobody joins them up.

The VDF AI Opportunity

Full coverage, with the evidence kept

Coverage

Test Everything, Not Twenty-Five Items

The population is the sample.

Every transaction, approval and record in scope is tested against the audit programme rather than a statistical extract, which changes what is findable: concentration, timing patterns and repeated exceptions by the same actor become visible instead of improbable.

  • Whole population tested against the programme
  • Concentration and timing patterns surfaced
  • Exceptions grouped by actor and by control
  • Coverage stated explicitly in the workpaper
Full
Test Coverage

Not an extract

TransactionsApprovalsExceptionsPatterns

Evidence

Keep What The Conclusion Rests On

Re-examinable in November.

Each exception retains the passage, record or field that establishes it along with the criteria version applied, so a finding challenged later is settled by reopening the evidence rather than by repeating the fieldwork from the beginning.

Retained
Test Evidence

Per exception

Source recordCriteria versionTest dateTester

Output

A Workpaper, Not A Spreadsheet

Drafted in your own format.

The output is a structured workpaper — objective, criteria, population, method, exceptions, conclusion left open — in the format your function already uses, so the auditor edits and signs rather than transcribing results into a template.

Structured
Workpaper

Conclusion left open

ObjectiveMethodExceptionsOpen conclusion
Run sequence

How the AI Internal Audit Agent runs a task

  1. STEP 01

    Take the programme as written

    The audit programme, control descriptions and delegation matrices are read as the criteria, so testing measures against what your function committed to examine rather than against a generic control library.

    Programme parsingCriteria extraction
  2. STEP 02

    Establish the population

    The complete set of items in scope is enumerated and reconciled to a control total before testing begins, because a test over an incomplete population produces a conclusion that is confident and unsupported.

    Population extractionCompleteness reconciliation
  3. STEP 03

    Test every item

    Each item is examined against the criteria, and the result is recorded as pass, exception or unable to test, with the third category kept distinct rather than being quietly counted as compliant.

    Criteria testingResult classification
  4. STEP 04

    Bind the evidence

    Every exception keeps the field, passage or document that establishes it together with the criteria version and the date, so the test can be reopened and checked long after the engagement has closed.

    Evidence citationVersion stamping
  5. STEP 05

    Draft and hand over

    Results are assembled into a workpaper with the conclusion deliberately left blank, patterns across actors and periods are reported separately, and the auditor decides what is a finding and how it is rated.

    Workpaper draftingPattern analysisAuditor handover
Integrations

Systems the AI Internal Audit Agent connects to

Scoped, per-tenant credentials Every call written to the audit log No data copied to a third party
Specification

Inputs, outputs and runtime

Ingests
Audit programme and criteriaPopulation under testApproval and delegation dataSupporting evidence documentsPrior engagement findings
Produces
Drafted workpaperException schedule with evidenceUnable-to-test listRecurring finding analysisCoverage statement
Triggered by
Engagement fieldwork startContinuous audit runFollow-up verification
Human oversight
The auditor forms every opinion and rating
Models
Open-weight LLMs you host — Llama, Qwen or Mistral class
Typical latency
Hours for a full population test
Deployment
On-premise or sovereign cloud with egress control
Data residency
Audit evidence stays inside your network
Where it pays back

Where the Internal Audit Agent pays back

Full-Population Control Testing

Test every item against the control criteria rather than a sample, and report exceptions with their evidence.

Approval Completeness Review

Establish which transactions lack the approvals the delegation matrix requires, and who authorised them instead.

Segregation Of Duties Testing

Find the cases where the same person initiated and approved, across the full period rather than in a sample.

Workpaper Preparation

Draft the structured workpaper for each test with population, method and exceptions filled in for review.

Recurring Finding Analysis

Identify controls that have failed across several engagements and business units over successive years.

Follow-Up Verification

Re-test whether prior findings were actually remediated rather than marked closed on a management assertion.

Comparison

AI Internal Audit Agent vs chatbots and SaaS copilots

Sampling exists because a person can read twenty-five items and not forty thousand, and every methodology built on top of it is an attempt to make that constraint respectable rather than a claim that it is ideal.

  Generic chatbot SaaS copilot VDF AI
Coverage What you paste A sample The whole population
Criteria Generic practice Standard library Your own audit programme
Evidence Not kept Summarised Source record per exception
Unable to test Treated as pass Omitted Reported as its own category
Recurring findings Invisible Per engagement Joined across engagements
Forms the opinion Freely Suggests a rating Never — the auditor does
Where evidence is read Vendor service Vendor tenancy Inside your own network
Controls

Governance and controls

Independence is the whole basis of the third line, so an audit agent has to be demonstrably incapable of altering what it tests and demonstrably unable to reach the conclusion on the auditor’s behalf.

IIA standardsSOX-style controlsISO 27001Internal audit charter

Read-only on tested records

Nothing under audit can be modified

Conclusion left blank

The workpaper ships without an opinion

Criteria version recorded

Each test states the programme used

Unable-to-test preserved

Missing evidence never counts as pass

Evidence retained

Source records kept for re-examination

Auditor attribution

A named auditor signs every workpaper

Evidence it leaves behind

Population reconciliation Per-item test result Exception evidence record Auditor sign-off trail
ROI snapshot

What changes after rollout

Complete Testing across the whole population
Repeatable Findings re-examinable from their evidence
Shorter Fieldwork time per audit engagement
Joined up Recurring control failures made visible
Audience

Who runs the AI Internal Audit Agent

Head of internal audit

Can tell the committee what proportion of a population was tested rather than describing a sampling approach, and can show which control failures have now recurred across three engagements in different units.

Audit senior

Arrives at fieldwork with the population reconciled and the exceptions already itemised with their evidence, which moves the engagement from document collection to the professional judgement it was scoped for.

Control owner in the business

Receives exceptions that quote the specific record and the criterion it failed, so the response is a factual correction rather than an argument about whether the auditor understood the process.

FAQ

Questions about the AI Internal Audit Agent

What is an AI internal audit agent?

It is an agent that performs internal audit fieldwork: testing the whole population against your audit programme rather than a sample, retaining the evidence behind every exception, and drafting structured workpapers for an auditor to review and sign.

How is an AI internal audit agent different from a generic chatbot?

A chatbot can describe audit practice in general. This agent applies your own audit programme to your own records and keeps the specific document that establishes each exception it raises.

Can an AI internal audit agent run on-premise on audit evidence data?

Yes. Audit evidence spans payroll, payments, contracts and personnel records, and an assurance function that exported it to a third party would be creating the very exposure it exists to test for.

What does an AI internal audit agent produce, and in what format?

Structured workpapers with objective, criteria, population, method and exceptions, each exception carrying its source record, plus a recurring-finding analysis and coverage statement.

Where does an AI internal audit agent fit in a governed AI programme?

It performs fieldwork, not judgement. The audit opinion, the risk rating and the decision to raise a finding remain the auditor’s, and the agent never alters a record it tests.

Does testing the whole population replace sampling methodology?

Where the population is machine-readable, yes — and that removes the sampling risk entirely rather than quantifying it. Where evidence exists only on paper or in systems with no accessible interface, sampling still applies and the workpaper says which parts of the population were tested in full and which were sampled. Claiming full coverage over a population you could not actually read would be worse than sampling honestly.

How is this different from the AI Expense Audit Agent?

Scope. The expense agent tests one document type against one policy at high volume, and its output is a reviewer queue. This agent works an audit engagement against an audit programme over any subject matter — payments, access, procurement, payroll, change control — and its output is a workpaper that has to survive external review. The expense agent is a control; this is assurance over controls.

Can it decide whether something is a finding?

No. It reports exceptions against criteria, which is a factual matter, and the auditor decides which exceptions constitute a finding, how it is rated and whether it is reportable. That distinction carries real weight: an exception can be immaterial, already known, or explained by a compensating control the test could not see, and all three are judgements about context rather than about the data.

What happens when evidence for an item simply does not exist?

It is recorded as unable to test, kept as its own category, and reported with a count. This matters more than it sounds. The common failure in automated testing is for missing evidence to fall through as a pass, which converts an absence of assurance into positive assurance — precisely the error most likely to be found later by someone else.

How does it relate to the cybersecurity agents?

The security analyst agent tests whether technical security controls are deployed across the estate, working from configuration and telemetry. This agent tests business process controls — approvals, segregation, authorisation limits, policy adherence — working from records and documents. An audit of access management would draw on both, with the security agent supplying the technical state and this one testing whether the governing process was followed.

Test the population, not a sample of it

See the AI Internal Audit Agent run a control test and draft the workpaper.