AI Agent for Internal Audit Fieldwork
Sampling was never a methodology, it was a concession to how long reading takes. This agent tests the full population against your audit programme, records the document behind each exception, and hands the auditor a drafted workpaper instead of a spreadsheet to start from.
What is an AI internal audit agent?
An AI internal audit agent is a governed software worker that carries out internal audit fieldwork. It tests a full population against an organisation’s own audit programme instead of a sample, retains the record establishing each exception, identifies findings that recur across engagements, and drafts structured workpapers for an auditor to review.
What it does
What it is not
A five percent sample and a finding nobody can reproduce
Internal audit is asked for assurance over a population it can only read a fraction of, so a clean sample becomes an opinion about the whole. Meanwhile the fieldwork that did happen is recorded as a conclusion rather than as evidence, and when the finding is challenged six months later, reconstructing why it was reached is its own project.
Sampling misses the pattern
Twenty-five items out of forty thousand cannot detect behaviour that concentrates in a small, deliberate corner of the population.
Evidence is summarised away
A workpaper records that approval was obtained without keeping the record that establishes it, so the test is not repeatable.
Fieldwork consumes the audit
Most of an engagement is pulling documents and reconciling lists, leaving the least time for the judgement the auditor is there to apply.
Findings recur without being seen
The same control fails in three consecutive audits of different business units and nobody joins them up.
Full coverage, with the evidence kept
Coverage
Test Everything, Not Twenty-Five Items
The population is the sample.
Every transaction, approval and record in scope is tested against the audit programme rather than a statistical extract, which changes what is findable: concentration, timing patterns and repeated exceptions by the same actor become visible instead of improbable.
- Whole population tested against the programme
- Concentration and timing patterns surfaced
- Exceptions grouped by actor and by control
- Coverage stated explicitly in the workpaper
Not an extract
Evidence
Keep What The Conclusion Rests On
Re-examinable in November.
Each exception retains the passage, record or field that establishes it along with the criteria version applied, so a finding challenged later is settled by reopening the evidence rather than by repeating the fieldwork from the beginning.
Per exception
Output
A Workpaper, Not A Spreadsheet
Drafted in your own format.
The output is a structured workpaper — objective, criteria, population, method, exceptions, conclusion left open — in the format your function already uses, so the auditor edits and signs rather than transcribing results into a template.
Conclusion left open
How the AI Internal Audit Agent runs a task
- STEP 01
Take the programme as written
The audit programme, control descriptions and delegation matrices are read as the criteria, so testing measures against what your function committed to examine rather than against a generic control library.
Programme parsingCriteria extraction - STEP 02
Establish the population
The complete set of items in scope is enumerated and reconciled to a control total before testing begins, because a test over an incomplete population produces a conclusion that is confident and unsupported.
Population extractionCompleteness reconciliation - STEP 03
Test every item
Each item is examined against the criteria, and the result is recorded as pass, exception or unable to test, with the third category kept distinct rather than being quietly counted as compliant.
Criteria testingResult classification - STEP 04
Bind the evidence
Every exception keeps the field, passage or document that establishes it together with the criteria version and the date, so the test can be reopened and checked long after the engagement has closed.
Evidence citationVersion stamping - STEP 05
Draft and hand over
Results are assembled into a workpaper with the conclusion deliberately left blank, patterns across actors and periods are reported separately, and the auditor decides what is a finding and how it is rated.
Workpaper draftingPattern analysisAuditor handover
Systems the AI Internal Audit Agent connects to
Records under test
Criteria and analysis
Inputs, outputs and runtime
- Ingests
- Audit programme and criteriaPopulation under testApproval and delegation dataSupporting evidence documentsPrior engagement findings
- Produces
- Drafted workpaperException schedule with evidenceUnable-to-test listRecurring finding analysisCoverage statement
- Triggered by
- Engagement fieldwork startContinuous audit runFollow-up verification
- Human oversight
- The auditor forms every opinion and rating
- Models
- Open-weight LLMs you host — Llama, Qwen or Mistral class
- Typical latency
- Hours for a full population test
- Deployment
- On-premise or sovereign cloud with egress control
- Data residency
- Audit evidence stays inside your network
Where the Internal Audit Agent pays back
Full-Population Control Testing
Test every item against the control criteria rather than a sample, and report exceptions with their evidence.
Approval Completeness Review
Establish which transactions lack the approvals the delegation matrix requires, and who authorised them instead.
Segregation Of Duties Testing
Find the cases where the same person initiated and approved, across the full period rather than in a sample.
Workpaper Preparation
Draft the structured workpaper for each test with population, method and exceptions filled in for review.
Recurring Finding Analysis
Identify controls that have failed across several engagements and business units over successive years.
Follow-Up Verification
Re-test whether prior findings were actually remediated rather than marked closed on a management assertion.
AI Internal Audit Agent vs chatbots and SaaS copilots
Sampling exists because a person can read twenty-five items and not forty thousand, and every methodology built on top of it is an attempt to make that constraint respectable rather than a claim that it is ideal.
| Generic chatbot | SaaS copilot | VDF AI | |
|---|---|---|---|
| Coverage | What you paste | A sample | The whole population |
| Criteria | Generic practice | Standard library | Your own audit programme |
| Evidence | Not kept | Summarised | Source record per exception |
| Unable to test | Treated as pass | Omitted | Reported as its own category |
| Recurring findings | Invisible | Per engagement | Joined across engagements |
| Forms the opinion | Freely | Suggests a rating | Never — the auditor does |
| Where evidence is read | Vendor service | Vendor tenancy | Inside your own network |
Governance and controls
Independence is the whole basis of the third line, so an audit agent has to be demonstrably incapable of altering what it tests and demonstrably unable to reach the conclusion on the auditor’s behalf.
Read-only on tested records
Nothing under audit can be modified
Conclusion left blank
The workpaper ships without an opinion
Criteria version recorded
Each test states the programme used
Unable-to-test preserved
Missing evidence never counts as pass
Evidence retained
Source records kept for re-examination
Auditor attribution
A named auditor signs every workpaper
Evidence it leaves behind
What changes after rollout
Who runs the AI Internal Audit Agent
Head of internal audit
Can tell the committee what proportion of a population was tested rather than describing a sampling approach, and can show which control failures have now recurred across three engagements in different units.
Audit senior
Arrives at fieldwork with the population reconciled and the exceptions already itemised with their evidence, which moves the engagement from document collection to the professional judgement it was scoped for.
Control owner in the business
Receives exceptions that quote the specific record and the criterion it failed, so the response is a factual correction rather than an argument about whether the auditor understood the process.
Questions about the AI Internal Audit Agent
What is an AI internal audit agent?
It is an agent that performs internal audit fieldwork: testing the whole population against your audit programme rather than a sample, retaining the evidence behind every exception, and drafting structured workpapers for an auditor to review and sign.
How is an AI internal audit agent different from a generic chatbot?
A chatbot can describe audit practice in general. This agent applies your own audit programme to your own records and keeps the specific document that establishes each exception it raises.
Can an AI internal audit agent run on-premise on audit evidence data?
Yes. Audit evidence spans payroll, payments, contracts and personnel records, and an assurance function that exported it to a third party would be creating the very exposure it exists to test for.
What does an AI internal audit agent produce, and in what format?
Structured workpapers with objective, criteria, population, method and exceptions, each exception carrying its source record, plus a recurring-finding analysis and coverage statement.
Where does an AI internal audit agent fit in a governed AI programme?
It performs fieldwork, not judgement. The audit opinion, the risk rating and the decision to raise a finding remain the auditor’s, and the agent never alters a record it tests.
Does testing the whole population replace sampling methodology?
Where the population is machine-readable, yes — and that removes the sampling risk entirely rather than quantifying it. Where evidence exists only on paper or in systems with no accessible interface, sampling still applies and the workpaper says which parts of the population were tested in full and which were sampled. Claiming full coverage over a population you could not actually read would be worse than sampling honestly.
How is this different from the AI Expense Audit Agent?
Scope. The expense agent tests one document type against one policy at high volume, and its output is a reviewer queue. This agent works an audit engagement against an audit programme over any subject matter — payments, access, procurement, payroll, change control — and its output is a workpaper that has to survive external review. The expense agent is a control; this is assurance over controls.
Can it decide whether something is a finding?
No. It reports exceptions against criteria, which is a factual matter, and the auditor decides which exceptions constitute a finding, how it is rated and whether it is reportable. That distinction carries real weight: an exception can be immaterial, already known, or explained by a compensating control the test could not see, and all three are judgements about context rather than about the data.
What happens when evidence for an item simply does not exist?
It is recorded as unable to test, kept as its own category, and reported with a count. This matters more than it sounds. The common failure in automated testing is for missing evidence to fall through as a pass, which converts an absence of assurance into positive assurance — precisely the error most likely to be found later by someone else.
How does it relate to the cybersecurity agents?
The security analyst agent tests whether technical security controls are deployed across the estate, working from configuration and telemetry. This agent tests business process controls — approvals, segregation, authorisation limits, policy adherence — working from records and documents. An audit of access management would draw on both, with the security agent supplying the technical state and this one testing whether the governing process was followed.
Test the population, not a sample of it
See the AI Internal Audit Agent run a control test and draft the workpaper.