AI in Auditing
Real uses, accuracy evidence, standards, and examples—AI in audit that keeps human responsibility intact.
Educational content — not assurance advice or a substitute for professional judgment. Auditing standards and firm tools change; recheck current AICPA, PCAOB, IAASB, and vendor sources before relying on a specific capability or requirement. Research accuracy figures on this page are tied to the cited study, model, dataset, and conditions—never a universal claim or a guarantee for any single engagement.
Why This Matters
AI is reshaping how audits are performed, but it does not change what an audit is or who is responsible for it. In practice, AI and automated tools help auditors analyze entire populations of transactions instead of samples, flag anomalies for follow-up, extract data from documents, and draft memos and communications. Deloitte's Omnia platform, for example, uses generative AI to perform initial documentation reviews, help navigate draft financial statements, extract and summarize data, and draft first-pass audit communications and accounting memos. But the auditing standards are explicit: automated tools do not replace the auditor's involvement, judgment, and professional skepticism, and they do not shift responsibility for the audit away from the humans on the engagement.
The central lesson mirrors the other modules in this series: AI can make audit procedures more powerful and efficient, but the auditor remains responsible for evaluating evidence, exercising skepticism, supervising the work, and documenting the basis for every conclusion. Two facts make this especially important in auditing. First, peer-reviewed evidence shows AI-based fraud-detection models are useful but imperfect, with accuracy that varies widely by method and dataset. Second—and counterintuitively—experimental research shows auditors do not always over-trust AI; a landmark study found they sometimes under-react to evidence when it comes from an AI system. Both over-reliance and under-reliance threaten audit quality, which is why calibrated human judgment is the skill that matters most.
Learning objective
By the end of this module, learners should be able to describe how AI is actually used in auditing, distinguish "auditing with AI" from "auditing the AI a client uses," explain what the auditing standards require when technology is involved, interpret research on AI fraud-detection accuracy and automation bias, and apply a governance framework that keeps human responsibility intact.
What "AI in Auditing" Actually Means
AI in auditing is the use of data analytics, machine learning, and generative AI to assist auditors in planning and performing an audit and in evaluating evidence. It is important to separate two very different ideas that are often conflated.
Auditing with AI
The audit firm uses AI tools to perform or support audit procedures—analyzing transactions, detecting anomalies, extracting data, or drafting documentation.
Auditing the AI
The client uses AI in its own financial reporting or controls, so the auditor must understand and test those systems as part of the audit.
One systematic review argues firms should explicitly separate these two activities in their methodologies, because they require different evidence and competencies, and the literature frequently blurs them. This page focuses primarily on auditing with AI, while noting where auditing the AI applies.
ADA / rules vs. generative AI
A second essential distinction is between two categories of tool. Audit data analytics (ADA) and rule-based automation apply defined logic to data—recalculating, matching, filtering, or aggregating—and are comparatively transparent. Generative AI produces text or content by prediction and can be fluent but wrong. As with the tax and bookkeeping modules, conflating deterministic analysis with generative prediction is the most dangerous mistake a learner can make.
How AI Enters the Audit
AI can appear in every phase of the audit, but each use routes back into the same evidence, documentation, and quality-control expectations that apply to any procedure.
| Audit phase | AI-assisted activity | What the standards still require |
|---|---|---|
| Risk assessment | Analyze full populations to identify unusual patterns and higher-risk areas | The auditor understands the entity's IT environment and evaluates whether data is relevant and reliable |
| Planning | Target testing toward high-risk segments | The engagement team designs and supervises procedures |
| Tests of controls | Test IT general controls and automated application controls over electronic information | Auditors evaluate the reliability of information used as evidence |
| Substantive procedures | Perform analytics across 100% of a population; identify items warranting investigation | Investigation must determine whether flagged items indicate misstatements or control deficiencies |
| Documentation | Draft memos, communications, and summaries | The file must reflect the information used, output produced, and reasoning supporting the conclusion |
| Review | Perform initial reviews of documentation for clarity | Supervision and review responsibilities are unchanged |
The right mental model is understand the data → analyze with technology → investigate exceptions → evaluate evidence → conclude and document. AI can dramatically expand the first three steps, but it cannot perform the evaluation and conclusion that connect evidence to the auditor's opinion.
Four Real Use Cases
Full-population analysis instead of sampling
Historically, auditors tested samples because examining every transaction was impractical. Automated tools and audit data analytics let auditors process large volumes of data from multiple sources and, in some cases, analyze a complete population rather than a sample. The AICPA notes this can produce more persuasive audit evidence as the volume of available information expands, and can help auditors identify specific items for further testing or understand characteristics across a population. This is powerful, but it is not automatic assurance. Using analytics does not necessarily eliminate the need for sampling, and analyzing a full population raises a practical challenge: how much work to perform on the exceptions and outliers the analysis surfaces. The PCAOB's amended standards are explicit that when auditors use technology to identify items meeting certain criteria, their investigation must determine whether those items—individually or in aggregate—indicate misstatements or control deficiencies.
Auditor's checkpoint: Full-population analysis changes what is examined, not the obligation to investigate exceptions and evaluate whether they represent misstatements.
Anomaly and fraud-risk detection
Machine-learning models can identify unusual patterns and anomalies that may warrant further investigation, and a large body of research tests their ability to flag potential financial-statement fraud. These tools help auditors look for items exhibiting specified attributes and organize data to reveal relationships. The research is genuinely promising but must be read carefully. A cross-country study using an ensemble model reported overall accuracy of 85.69%, correctly classifying about 82% of manipulated and about 90% of non-manipulated firm-years—yet when applied retrospectively to the Wirecard fraud, it flagged only 7 of 17 firm-years. That gap between aggregate accuracy and catching a specific real fraud is the key lesson: a strong average is not a guarantee for any single engagement.
Auditor's checkpoint: A model's flag is a starting point for professional investigation, not a conclusion about fraud.
Data and document extraction
Generative AI can extract and summarize information across documents, help auditors navigate draft financial statements, and answer questions about statement content to streamline tie-out procedures. This reduces manual effort in gathering and organizing evidence. But extraction is not evaluation. The PCAOB emphasizes that auditors must evaluate the reliability of electronic information used as evidence; if data is generated by a client system, the auditor must consider the IT controls that produced it. Reading a number correctly does not establish that the number is reliable.
Auditor's checkpoint: Verify the reliability of extracted information and its source, not just its accuracy of transcription.
Drafting and documentation support
Firms use generative AI to draft first-pass audit communications, accounting memos, and to perform initial documentation reviews. According to the PCAOB's 2024 outreach, the integration of generative AI in audits at the firms staff spoke with appeared focused primarily on administrative and research activities rather than forming audit conclusions. This is the lowest-risk category only if the output is treated as a draft. A tool that drafts a memo does not shrink the auditor's duty to supervise the work, evaluate the evidence, and document what was done and why.
Auditor's checkpoint: AI-drafted content is a starting draft that the responsible auditor must review, correct, and own.
What the Research Actually Shows
The evidence base supports a nuanced conclusion: AI meaningfully enhances detection and efficiency, but its accuracy is variable, its outputs require investigation, and human reliance on it is not always well-calibrated.
Fraud-detection accuracy is useful but imperfect and highly variable
Peer-reviewed studies report a wide range of accuracy depending on the algorithm, features, and dataset. One review of prior literature found an average detection accuracy across studies of about 78%. Individual studies report higher figures under specific conditions: an ensemble model reaching 85.69% overall accuracy on a global dataset; Random Forest reaching about 84% accuracy with an AUC around 0.92 on a Vietnamese dataset; and gradient-boosting models reaching roughly 94% accuracy in another comparative study.
| Study context | Reported performance |
|---|---|
| Global cross-country ensemble (gradient boosting + KNN) | 85.69% overall accuracy; 82.03% of fraud and 89.88% of non-fraud correctly classified; flagged 7 of 17 Wirecard firm-years |
| Literature average across analyzed studies | ~78.37% average detection accuracy |
| Vietnamese listed firms (Random Forest) | ~84% accuracy; AUC ~0.92; recall ~78% |
| Comparative study (Gradient Boosting Machine) | 94% accuracy; AUC-ROC 0.96 |
Study-bound figures: These numbers come from different countries, datasets, class balances, and fraud definitions. Higher accuracy is often achieved on balanced or oversampled datasets that do not reflect the extreme rarity of fraud in the real world. A single headline accuracy figure means little without its dataset and conditions—and is never a guarantee for any single engagement. Models are valuable for surfacing risk but cannot, on their own, reliably conclude that fraud exists.
Auditors can under-rely on AI, not just over-rely
The most important behavioral finding for auditing runs against common intuition. In a controlled experiment with 170 audit seniors published in the Journal of Accounting Research, Commerford, Dennis, Joe, and Ulla (2022) found that auditors who received evidence contradicting management's complex estimate proposed smaller adjustments when that evidence came from their firm's AI system than when identical evidence came from a human specialist—an effect the authors attribute to "algorithm aversion." The effect was especially pronounced when management's estimate used relatively objective inputs, and the authors warn this susceptibility "could prove costly for the profession and financial statement users."
Subsequent work shows reliance is contingent on task complexity and uncertainty, and that auditors may rely more on adaptive "learning" algorithms than on static ones in high-uncertainty settings. A separate survey-based study found that AI is widely used and seen as useful for fraud detection, but did not find a significant link between AI usage and stronger professional skepticism—concluding that AI improves speed and analytical support but does not automatically improve skeptical judgment.
Both failure modes threaten audit quality
The literature therefore documents two opposing risks. Broader research on automation bias—the tendency to uncritically accept algorithmic output—warns that auditors can over-rely on AI, particularly when outputs are presented with high confidence or align with their initial assessments. At the same time, the Commerford et al. experiment shows auditors can under-rely on AI that contradicts management. The practical implication is that professional skepticism must become a calibration practice: auditors should neither accept AI outputs uncritically nor dismiss them outright, but weigh them against risk, context, and other evidence.
A note on evidence quality
Learners should also know the evidence base itself has limits. One appraisal found that much of the AI-in-audit literature makes prescriptive claims without archival studies linking AI use to measured audit outcomes such as opinion accuracy, audit fees, or inspection results. In other words, strong claims that "AI improves audit quality" outrun the peer-reviewed evidence currently available, which is a reason for measured, well-documented adoption rather than sweeping conclusions.
What the Auditing Standards Require
The standards do not create a separate rulebook for AI. Existing requirements on evidence, supervision, documentation, and quality control apply to any work an AI tool touches.
AICPA (private-company audits under GAAS)
In October 2025 the AICPA published Technical Questions and Answers Section 8400, Use of Technology in an Audit of Financial Statements. Key points include:
- GAAS neither requires nor precludes the use of automated tools and techniques.
- Automated tools may be used in risk assessment, analytical procedures, and tests of controls, and can help process large volumes of complex data and identify items for further investigation.
- Critically, automated tools "don't replace the need for auditor involvement, judgment, and skepticism in all audit phases," and auditors must understand the entity's IT environment and evaluate whether electronic data is relevant and reliable.
These Q&As build on SAS No. 142, Audit Evidence, which recognized automated tools and techniques—including audit data analytics—and SAS No. 145 on risk assessment, for which the AICPA issued a practice aid on using technology.
PCAOB (public-company audits)
In June 2024 the PCAOB adopted amendments to AS 1105, Audit Evidence, and AS 2301, The Auditor's Responses to the Risks of Material Misstatement, addressing technology-assisted analysis of electronic information. The amendments, effective for audits of fiscal years beginning on or after December 15, 2025, clarify that auditors must evaluate the reliability of electronic information used as evidence, must test relevant IT general controls and automated application controls where they rely on controls, must achieve each objective when a procedure serves multiple purposes, and must investigate whether items identified by technology indicate misstatements or control deficiencies.
PCAOB Spotlight note: The PCAOB's July 2024 staff Spotlight on generative AI is not a rule, standard, or safe harbor—it represents staff views based on outreach to firms. It reports that current GenAI use is concentrated in administrative and research tasks, that some firms are keeping GenAI out of audit procedures for now over privacy and reliability concerns, and that AI-assisted procedures must be documented like any other procedure. The staff emphasize that supervision under AS 1201 and documentation under AS 1215 apply to AI-assisted work exactly as they do to any other audit procedure, and that AI does not shift responsibility for due professional care or professional skepticism.
Quality control over the tools themselves
Under the PCAOB's quality-control standard QC 1000, approving and monitoring AI tools is part of a firm's quality-control system rather than a purely IT decision, and the PCAOB is already asking firms how they use these tools during inspections. Before adopting a tool, a firm should be able to explain whether the tool can be audited end-to-end, whether a third-party model's behavior can be explained, and whether bias and auditability were considered for the specific procedures the tool supports.
IAASB (international standards)
Internationally, the ISAs "do not prohibit, nor stimulate" the use of data analytics; they acknowledge the auditor's use of technology through computer-assisted audit techniques. The IAASB has identified issues auditors must consider when applying data analytics, including IT controls at both the entity and the auditor, the completeness and reliability of information produced by the entity, how much work to perform on identified exceptions, documentation, and the reliability of third-party analytics tools.
Worked Example: Auditing ABC Coffee Shop's Revenue
The following example is hypothetical and illustrates verification logic, not a specific software screen. An audit team is testing revenue for ABC Coffee Shop, a company with tens of thousands of point-of-sale transactions. Rather than sampling, the team uses audit data analytics to examine the full population and applies an anomaly-detection model to flag unusual entries.
| Situation | AI behavior | Correct auditor response |
|---|---|---|
| Full-population analysis of sales | Tool recalculates and totals every transaction | Confirm the data is complete and reliable, and that IT controls over the POS data were considered |
| Anomaly detection | Model flags 240 unusual transactions | Investigate whether flagged items indicate misstatements or control deficiencies—individually and in aggregate |
| Exceptions surfaced | Analysis reveals many outliers | Decide and document how much work each exception warrants; do not ignore outliers |
| GenAI drafts the revenue memo | Tool produces a first-draft workpaper | Review, correct, and take responsibility for the memo; document the reasoning |
| Model contradicts management's estimate | AI evidence suggests a needed adjustment | Weigh it as seriously as human evidence—guard against under-reacting to AI-sourced contradiction |
Why a clean analytics run is not a conclusion
The analytics can confirm that every transaction was captured and totaled and can surface anomalies far faster than manual review. It cannot conclude whether the anomalies represent errors, fraud, or legitimate business activity—that determination requires professional investigation and skepticism. And if the model's evidence contradicts management, the research warns the team may unconsciously discount it precisely because it came from a machine.
A safer audit sequence
Understand the data and its controls
Establish that the population is complete and reliable before analyzing it.
Analyze the full population
Use ADA to identify risks and target testing.
Investigate every exception appropriately
Determine and document the work performed on outliers.
Evaluate the evidence with calibrated skepticism
Weigh AI-sourced evidence neither too lightly nor too heavily.
Supervise and review
Ensure the engagement team designed, performed, and reviewed the work.
Document the file
Record the information used, the output produced, and the reasoning supporting the conclusion.
Risks Auditors Must Control
Automation bias (over-reliance)
Auditors can uncritically accept AI output, especially when it is confident or confirms their expectations, eroding the professional skepticism foundational to audit quality.
Control: Treat AI output as evidence to be evaluated, not a conclusion; require documented investigation of flagged items.
Algorithm aversion (under-reliance)
Conversely, auditors may discount AI evidence that contradicts management, proposing smaller adjustments than warranted.
Control: Apply skepticism as calibration—weigh AI-sourced contradictory evidence as seriously as human-sourced evidence.
Unreliable data and "garbage in"
Analytics on incomplete or manipulated data produce confident but wrong results; the reliability of electronic evidence is the auditor's responsibility.
Control: Evaluate data completeness and reliability, and test relevant IT general and application controls before relying on results.
Opaque or unexplainable models
A model whose behavior cannot be explained undermines the auditor's ability to defend the professional judgment behind a conclusion.
Control: Before adoption, confirm the tool can be audited and, for third-party models, that behavior can be explained for the specific procedures it supports.
Generative AI errors in drafting and research
Generative AI can produce fluent but inaccurate memos or research; the file must still reflect correct, supported reasoning.
Control: Treat GenAI output as a draft; the responsible auditor reviews, verifies, and owns it, and documents the work.
Weak documentation and quality control
AI-assisted procedures that are not documented like any other procedure fail supervision and documentation requirements.
Control: Document the information used, output produced, and reasoning; govern tool approval and monitoring under the firm's quality-control system.
A Control Framework for AI in an Audit Practice
The structure below mirrors the profession's own emphasis: govern the tools, understand the data, investigate results, and document the basis for conclusions.
| Control area | Practical question | Minimum evidence |
|---|---|---|
| Tool governance | Is the tool approved and monitored under the firm's quality-control system? | QC 1000 approval and monitoring record |
| Explainability | Can the tool be audited end-to-end and its behavior explained? | Tool evaluation and vendor documentation |
| Data reliability | Is the electronic information complete and reliable? | Tests of IT general and application controls |
| Exception handling | How much work is performed on flagged outliers? | Documented investigation of exceptions |
| Skepticism calibration | Is AI evidence weighed neither too lightly nor too heavily? | Evaluation notes and adjustments |
| Supervision | Did the engagement team design, perform, and review the work? | Supervision record under AS 1201 |
| Documentation | Does the file reflect information used, output, and reasoning? | Workpapers under AS 1215 |
| Auditing client AI | Where the client uses AI in reporting, is it understood and tested? | Understanding of client systems and controls |
Suggested operating policy
- AI and analytics may analyze populations, flag anomalies, extract data, and draft documentation.
- AI output is evidence to be evaluated, never a conclusion, and never a substitute for skepticism.
- Every exception surfaced by technology is investigated and the work is documented.
- Data reliability and relevant IT controls are evaluated before results are relied upon.
- Every AI tool is approved and monitored under the firm's quality-control system.
- The engagement team supervises, reviews, and documents all AI-assisted work.
- The auditor remains responsible for the opinion.
Thresholds and specifics should reflect the engagement's risk, the tools used, and applicable standards (GAAS, PCAOB, or ISAs).
The Changing Auditor Role
The evidence supports a shift in task composition, not the disappearance of the auditor. AI can analyze full populations, flag anomalies, extract data, and draft documentation. The auditor's work moves toward understanding data and systems, investigating exceptions, exercising calibrated skepticism, supervising the engagement, and taking responsibility for the opinion.
That shift raises the value of four abilities:
Data and technology literacy
Understanding what a tool does, where it fails, and whether its data is reliable.
Calibrated professional skepticism
Neither over-trusting nor reflexively dismissing AI output.
Investigative judgment
Turning a flagged anomaly into a supported conclusion about misstatement.
Documentation discipline
Recording the information used, the output produced, and the reasoning behind every conclusion.
The strongest future auditor is not the one who reviews samples fastest. It is the one who can supervise powerful tools, recognize when they are wrong in either direction, investigate what they surface, and stand behind the audit opinion.
Practice: Evaluate the AI Behavior
For each scenario, decide whether to Rely, Investigate, or Challenge, then reveal the best response.
Scenario 1
An analytics tool confirms that 100% of recorded sales transactions were captured and correctly totaled.
Scenario 2
An anomaly-detection model flags 300 transactions as unusual.
Scenario 3
The firm's AI system produces evidence contradicting management's complex estimate, suggesting a material adjustment.
Scenario 4
Generative AI drafts an accounting memo that reads well and cites a conclusion.
Scenario 5
A firm wants to deploy a third-party AI model whose internal logic the vendor will not explain.
Score: 0/5 correct (0 reviewed)
Key Takeaways
- AI in auditing is real and supports full-population analysis, anomaly detection, data extraction, and documentation drafting inside firm platforms such as Deloitte Omnia.
- "Auditing with AI" and "auditing the AI a client uses" are different activities requiring different evidence.
- Fraud-detection models are useful but imperfect—accuracy averages roughly 78% across studies and varies widely, and strong aggregate accuracy does not guarantee catching a specific fraud.
- Auditor reliance on AI is not always well-calibrated: research documents both automation bias (over-reliance) and algorithm aversion (under-reliance), so skepticism must be a calibration practice.
- The standards do not create a separate AI rulebook: evidence reliability, supervision (AS 1201), documentation (AS 1215), and quality control (QC 1000) all apply to AI-assisted work.
- The auditor remains responsible for evaluating evidence, investigating exceptions, and issuing the opinion.
The goal is not auditing without auditors. The goal is auditing in which auditors spend less time on manual sampling and more time investigating what powerful tools surface, weighing evidence with calibrated skepticism, and standing behind a defensible opinion.
Knowledge Check
Five questions on auditing with vs. the AI, research findings, algorithm aversion, standards responsibility, and responsible practice.
Question 1: What is the difference between "auditing with AI" and "auditing the AI"?
Question 2: What does the peer-reviewed research show about AI fraud-detection accuracy?
Question 3: What did Commerford et al. (2022) find about auditor reliance on AI?
Question 4: Under the auditing standards, what happens to auditor responsibility when AI is used?
Question 5: Which statement best describes responsible AI-assisted auditing?
Sources & Further Reading
Selected primary, research, and practice sources used in this module. Auditing standards and firm tools change; recheck official AICPA, PCAOB, IAASB, and vendor docs when evaluating a tool or procedure. Research accuracy figures remain tied to their study conditions.
- Thomson Reuters — AICPA Issues Technical Q&As on Technology Use in Financial Statement Audits (TQA 8400)
- PCAOB — Updates Standards to Clarify Auditor Responsibilities When Using Technology-Assisted Analysis
- The Leveraged Years — PCAOB Staff Spotlight: How Auditors Use Generative AI (staff views, not a rule)
- Deloitte — Expands AI Capabilities in Omnia
- Commerford et al. (2022) — Man Versus Machine: Complex Estimates and Auditor Reliance on Artificial Intelligence (Journal of Accounting Research / RePEc)
- MDPI — Using Machine Learning to Detect Financial Statement Fraud: A Cross-Country Analysis Applied to Wirecard AG
- Fieldguide — PCAOB AI Oversight: Where the Board Stands on AI in Audit
- Baker Tilly — Beyond the Algorithm: What AI Adoption in Public Company Finance Means for Governance
Ready to Practice?
Take the investigation and calibration habits from this lesson into the AI Practice Arena—judgment first, automation second.
Open AI Practice ArenaWhat's Next?
Next in the practice catalog is AI in Financial Analysis (coming soon)—or return to the AI hub to revisit Foundations and related practice topics.
AI in Financial Analysis
Next in the practice catalog (coming soon)
AI in Accounting Hub
Browse all AI pillar topics