Skip to main content
Concept #128

AI in Auditing

Real uses, accuracy evidence, standards, and examples—AI in audit that keeps human responsibility intact.

Educational content — not assurance advice or a substitute for professional judgment. Auditing standards and firm tools change; recheck current AICPA, PCAOB, IAASB, and vendor sources before relying on a specific capability or requirement. Research accuracy figures on this page are tied to the cited study, model, dataset, and conditions—never a universal claim or a guarantee for any single engagement.

Why This Matters

AI is reshaping how audits are performed, but it does not change what an audit is or who is responsible for it. In practice, AI and automated tools help auditors analyze entire populations of transactions instead of samples, flag anomalies for follow-up, extract data from documents, and draft memos and communications. Deloitte's Omnia platform, for example, uses generative AI to perform initial documentation reviews, help navigate draft financial statements, extract and summarize data, and draft first-pass audit communications and accounting memos. But the auditing standards are explicit: automated tools do not replace the auditor's involvement, judgment, and professional skepticism, and they do not shift responsibility for the audit away from the humans on the engagement.

The central lesson mirrors the other modules in this series: AI can make audit procedures more powerful and efficient, but the auditor remains responsible for evaluating evidence, exercising skepticism, supervising the work, and documenting the basis for every conclusion. Two facts make this especially important in auditing. First, peer-reviewed evidence shows AI-based fraud-detection models are useful but imperfect, with accuracy that varies widely by method and dataset. Second—and counterintuitively—experimental research shows auditors do not always over-trust AI; a landmark study found they sometimes under-react to evidence when it comes from an AI system. Both over-reliance and under-reliance threaten audit quality, which is why calibrated human judgment is the skill that matters most.

Learning objective

By the end of this module, learners should be able to describe how AI is actually used in auditing, distinguish "auditing with AI" from "auditing the AI a client uses," explain what the auditing standards require when technology is involved, interpret research on AI fraud-detection accuracy and automation bias, and apply a governance framework that keeps human responsibility intact.

What "AI in Auditing" Actually Means

AI in auditing is the use of data analytics, machine learning, and generative AI to assist auditors in planning and performing an audit and in evaluating evidence. It is important to separate two very different ideas that are often conflated.

Auditing with AI

The audit firm uses AI tools to perform or support audit procedures—analyzing transactions, detecting anomalies, extracting data, or drafting documentation.

Auditing the AI

The client uses AI in its own financial reporting or controls, so the auditor must understand and test those systems as part of the audit.

One systematic review argues firms should explicitly separate these two activities in their methodologies, because they require different evidence and competencies, and the literature frequently blurs them. This page focuses primarily on auditing with AI, while noting where auditing the AI applies.

ADA / rules vs. generative AI

A second essential distinction is between two categories of tool. Audit data analytics (ADA) and rule-based automation apply defined logic to data—recalculating, matching, filtering, or aggregating—and are comparatively transparent. Generative AI produces text or content by prediction and can be fluent but wrong. As with the tax and bookkeeping modules, conflating deterministic analysis with generative prediction is the most dangerous mistake a learner can make.

How AI Enters the Audit

AI can appear in every phase of the audit, but each use routes back into the same evidence, documentation, and quality-control expectations that apply to any procedure.

Audit phaseAI-assisted activityWhat the standards still require
Risk assessmentAnalyze full populations to identify unusual patterns and higher-risk areasThe auditor understands the entity's IT environment and evaluates whether data is relevant and reliable
PlanningTarget testing toward high-risk segmentsThe engagement team designs and supervises procedures
Tests of controlsTest IT general controls and automated application controls over electronic informationAuditors evaluate the reliability of information used as evidence
Substantive proceduresPerform analytics across 100% of a population; identify items warranting investigationInvestigation must determine whether flagged items indicate misstatements or control deficiencies
DocumentationDraft memos, communications, and summariesThe file must reflect the information used, output produced, and reasoning supporting the conclusion
ReviewPerform initial reviews of documentation for claritySupervision and review responsibilities are unchanged

The right mental model is understand the data → analyze with technology → investigate exceptions → evaluate evidence → conclude and document. AI can dramatically expand the first three steps, but it cannot perform the evaluation and conclusion that connect evidence to the auditor's opinion.

Four Real Use Cases

1

Full-population analysis instead of sampling

Historically, auditors tested samples because examining every transaction was impractical. Automated tools and audit data analytics let auditors process large volumes of data from multiple sources and, in some cases, analyze a complete population rather than a sample. The AICPA notes this can produce more persuasive audit evidence as the volume of available information expands, and can help auditors identify specific items for further testing or understand characteristics across a population. This is powerful, but it is not automatic assurance. Using analytics does not necessarily eliminate the need for sampling, and analyzing a full population raises a practical challenge: how much work to perform on the exceptions and outliers the analysis surfaces. The PCAOB's amended standards are explicit that when auditors use technology to identify items meeting certain criteria, their investigation must determine whether those items—individually or in aggregate—indicate misstatements or control deficiencies.

Auditor's checkpoint: Full-population analysis changes what is examined, not the obligation to investigate exceptions and evaluate whether they represent misstatements.

2

Anomaly and fraud-risk detection

Machine-learning models can identify unusual patterns and anomalies that may warrant further investigation, and a large body of research tests their ability to flag potential financial-statement fraud. These tools help auditors look for items exhibiting specified attributes and organize data to reveal relationships. The research is genuinely promising but must be read carefully. A cross-country study using an ensemble model reported overall accuracy of 85.69%, correctly classifying about 82% of manipulated and about 90% of non-manipulated firm-years—yet when applied retrospectively to the Wirecard fraud, it flagged only 7 of 17 firm-years. That gap between aggregate accuracy and catching a specific real fraud is the key lesson: a strong average is not a guarantee for any single engagement.

Auditor's checkpoint: A model's flag is a starting point for professional investigation, not a conclusion about fraud.

3

Data and document extraction

Generative AI can extract and summarize information across documents, help auditors navigate draft financial statements, and answer questions about statement content to streamline tie-out procedures. This reduces manual effort in gathering and organizing evidence. But extraction is not evaluation. The PCAOB emphasizes that auditors must evaluate the reliability of electronic information used as evidence; if data is generated by a client system, the auditor must consider the IT controls that produced it. Reading a number correctly does not establish that the number is reliable.

Auditor's checkpoint: Verify the reliability of extracted information and its source, not just its accuracy of transcription.

4

Drafting and documentation support

Firms use generative AI to draft first-pass audit communications, accounting memos, and to perform initial documentation reviews. According to the PCAOB's 2024 outreach, the integration of generative AI in audits at the firms staff spoke with appeared focused primarily on administrative and research activities rather than forming audit conclusions. This is the lowest-risk category only if the output is treated as a draft. A tool that drafts a memo does not shrink the auditor's duty to supervise the work, evaluate the evidence, and document what was done and why.

Auditor's checkpoint: AI-drafted content is a starting draft that the responsible auditor must review, correct, and own.

What the Research Actually Shows

The evidence base supports a nuanced conclusion: AI meaningfully enhances detection and efficiency, but its accuracy is variable, its outputs require investigation, and human reliance on it is not always well-calibrated.

Fraud-detection accuracy is useful but imperfect and highly variable

Peer-reviewed studies report a wide range of accuracy depending on the algorithm, features, and dataset. One review of prior literature found an average detection accuracy across studies of about 78%. Individual studies report higher figures under specific conditions: an ensemble model reaching 85.69% overall accuracy on a global dataset; Random Forest reaching about 84% accuracy with an AUC around 0.92 on a Vietnamese dataset; and gradient-boosting models reaching roughly 94% accuracy in another comparative study.

Study contextReported performance
Global cross-country ensemble (gradient boosting + KNN)85.69% overall accuracy; 82.03% of fraud and 89.88% of non-fraud correctly classified; flagged 7 of 17 Wirecard firm-years
Literature average across analyzed studies~78.37% average detection accuracy
Vietnamese listed firms (Random Forest)~84% accuracy; AUC ~0.92; recall ~78%
Comparative study (Gradient Boosting Machine)94% accuracy; AUC-ROC 0.96

Study-bound figures: These numbers come from different countries, datasets, class balances, and fraud definitions. Higher accuracy is often achieved on balanced or oversampled datasets that do not reflect the extreme rarity of fraud in the real world. A single headline accuracy figure means little without its dataset and conditions—and is never a guarantee for any single engagement. Models are valuable for surfacing risk but cannot, on their own, reliably conclude that fraud exists.

Auditors can under-rely on AI, not just over-rely

The most important behavioral finding for auditing runs against common intuition. In a controlled experiment with 170 audit seniors published in the Journal of Accounting Research, Commerford, Dennis, Joe, and Ulla (2022) found that auditors who received evidence contradicting management's complex estimate proposed smaller adjustments when that evidence came from their firm's AI system than when identical evidence came from a human specialist—an effect the authors attribute to "algorithm aversion." The effect was especially pronounced when management's estimate used relatively objective inputs, and the authors warn this susceptibility "could prove costly for the profession and financial statement users."

Subsequent work shows reliance is contingent on task complexity and uncertainty, and that auditors may rely more on adaptive "learning" algorithms than on static ones in high-uncertainty settings. A separate survey-based study found that AI is widely used and seen as useful for fraud detection, but did not find a significant link between AI usage and stronger professional skepticism—concluding that AI improves speed and analytical support but does not automatically improve skeptical judgment.

Both failure modes threaten audit quality

The literature therefore documents two opposing risks. Broader research on automation bias—the tendency to uncritically accept algorithmic output—warns that auditors can over-rely on AI, particularly when outputs are presented with high confidence or align with their initial assessments. At the same time, the Commerford et al. experiment shows auditors can under-rely on AI that contradicts management. The practical implication is that professional skepticism must become a calibration practice: auditors should neither accept AI outputs uncritically nor dismiss them outright, but weigh them against risk, context, and other evidence.

A note on evidence quality

Learners should also know the evidence base itself has limits. One appraisal found that much of the AI-in-audit literature makes prescriptive claims without archival studies linking AI use to measured audit outcomes such as opinion accuracy, audit fees, or inspection results. In other words, strong claims that "AI improves audit quality" outrun the peer-reviewed evidence currently available, which is a reason for measured, well-documented adoption rather than sweeping conclusions.

What the Auditing Standards Require

The standards do not create a separate rulebook for AI. Existing requirements on evidence, supervision, documentation, and quality control apply to any work an AI tool touches.

AICPA (private-company audits under GAAS)

In October 2025 the AICPA published Technical Questions and Answers Section 8400, Use of Technology in an Audit of Financial Statements. Key points include:

  • GAAS neither requires nor precludes the use of automated tools and techniques.
  • Automated tools may be used in risk assessment, analytical procedures, and tests of controls, and can help process large volumes of complex data and identify items for further investigation.
  • Critically, automated tools "don't replace the need for auditor involvement, judgment, and skepticism in all audit phases," and auditors must understand the entity's IT environment and evaluate whether electronic data is relevant and reliable.

These Q&As build on SAS No. 142, Audit Evidence, which recognized automated tools and techniques—including audit data analytics—and SAS No. 145 on risk assessment, for which the AICPA issued a practice aid on using technology.

PCAOB (public-company audits)

In June 2024 the PCAOB adopted amendments to AS 1105, Audit Evidence, and AS 2301, The Auditor's Responses to the Risks of Material Misstatement, addressing technology-assisted analysis of electronic information. The amendments, effective for audits of fiscal years beginning on or after December 15, 2025, clarify that auditors must evaluate the reliability of electronic information used as evidence, must test relevant IT general controls and automated application controls where they rely on controls, must achieve each objective when a procedure serves multiple purposes, and must investigate whether items identified by technology indicate misstatements or control deficiencies.

PCAOB Spotlight note: The PCAOB's July 2024 staff Spotlight on generative AI is not a rule, standard, or safe harbor—it represents staff views based on outreach to firms. It reports that current GenAI use is concentrated in administrative and research tasks, that some firms are keeping GenAI out of audit procedures for now over privacy and reliability concerns, and that AI-assisted procedures must be documented like any other procedure. The staff emphasize that supervision under AS 1201 and documentation under AS 1215 apply to AI-assisted work exactly as they do to any other audit procedure, and that AI does not shift responsibility for due professional care or professional skepticism.

Quality control over the tools themselves

Under the PCAOB's quality-control standard QC 1000, approving and monitoring AI tools is part of a firm's quality-control system rather than a purely IT decision, and the PCAOB is already asking firms how they use these tools during inspections. Before adopting a tool, a firm should be able to explain whether the tool can be audited end-to-end, whether a third-party model's behavior can be explained, and whether bias and auditability were considered for the specific procedures the tool supports.

IAASB (international standards)

Internationally, the ISAs "do not prohibit, nor stimulate" the use of data analytics; they acknowledge the auditor's use of technology through computer-assisted audit techniques. The IAASB has identified issues auditors must consider when applying data analytics, including IT controls at both the entity and the auditor, the completeness and reliability of information produced by the entity, how much work to perform on identified exceptions, documentation, and the reliability of third-party analytics tools.

Worked Example: Auditing ABC Coffee Shop's Revenue

The following example is hypothetical and illustrates verification logic, not a specific software screen. An audit team is testing revenue for ABC Coffee Shop, a company with tens of thousands of point-of-sale transactions. Rather than sampling, the team uses audit data analytics to examine the full population and applies an anomaly-detection model to flag unusual entries.

SituationAI behaviorCorrect auditor response
Full-population analysis of salesTool recalculates and totals every transactionConfirm the data is complete and reliable, and that IT controls over the POS data were considered
Anomaly detectionModel flags 240 unusual transactionsInvestigate whether flagged items indicate misstatements or control deficiencies—individually and in aggregate
Exceptions surfacedAnalysis reveals many outliersDecide and document how much work each exception warrants; do not ignore outliers
GenAI drafts the revenue memoTool produces a first-draft workpaperReview, correct, and take responsibility for the memo; document the reasoning
Model contradicts management's estimateAI evidence suggests a needed adjustmentWeigh it as seriously as human evidence—guard against under-reacting to AI-sourced contradiction

Why a clean analytics run is not a conclusion

The analytics can confirm that every transaction was captured and totaled and can surface anomalies far faster than manual review. It cannot conclude whether the anomalies represent errors, fraud, or legitimate business activity—that determination requires professional investigation and skepticism. And if the model's evidence contradicts management, the research warns the team may unconsciously discount it precisely because it came from a machine.

A safer audit sequence

1

Understand the data and its controls

Establish that the population is complete and reliable before analyzing it.

2

Analyze the full population

Use ADA to identify risks and target testing.

3

Investigate every exception appropriately

Determine and document the work performed on outliers.

4

Evaluate the evidence with calibrated skepticism

Weigh AI-sourced evidence neither too lightly nor too heavily.

5

Supervise and review

Ensure the engagement team designed, performed, and reviewed the work.

6

Document the file

Record the information used, the output produced, and the reasoning supporting the conclusion.

Risks Auditors Must Control

Automation bias (over-reliance)

Auditors can uncritically accept AI output, especially when it is confident or confirms their expectations, eroding the professional skepticism foundational to audit quality.

Control: Treat AI output as evidence to be evaluated, not a conclusion; require documented investigation of flagged items.

Algorithm aversion (under-reliance)

Conversely, auditors may discount AI evidence that contradicts management, proposing smaller adjustments than warranted.

Control: Apply skepticism as calibration—weigh AI-sourced contradictory evidence as seriously as human-sourced evidence.

Unreliable data and "garbage in"

Analytics on incomplete or manipulated data produce confident but wrong results; the reliability of electronic evidence is the auditor's responsibility.

Control: Evaluate data completeness and reliability, and test relevant IT general and application controls before relying on results.

Opaque or unexplainable models

A model whose behavior cannot be explained undermines the auditor's ability to defend the professional judgment behind a conclusion.

Control: Before adoption, confirm the tool can be audited and, for third-party models, that behavior can be explained for the specific procedures it supports.

Generative AI errors in drafting and research

Generative AI can produce fluent but inaccurate memos or research; the file must still reflect correct, supported reasoning.

Control: Treat GenAI output as a draft; the responsible auditor reviews, verifies, and owns it, and documents the work.

Weak documentation and quality control

AI-assisted procedures that are not documented like any other procedure fail supervision and documentation requirements.

Control: Document the information used, output produced, and reasoning; govern tool approval and monitoring under the firm's quality-control system.

A Control Framework for AI in an Audit Practice

The structure below mirrors the profession's own emphasis: govern the tools, understand the data, investigate results, and document the basis for conclusions.

Control areaPractical questionMinimum evidence
Tool governanceIs the tool approved and monitored under the firm's quality-control system?QC 1000 approval and monitoring record
ExplainabilityCan the tool be audited end-to-end and its behavior explained?Tool evaluation and vendor documentation
Data reliabilityIs the electronic information complete and reliable?Tests of IT general and application controls
Exception handlingHow much work is performed on flagged outliers?Documented investigation of exceptions
Skepticism calibrationIs AI evidence weighed neither too lightly nor too heavily?Evaluation notes and adjustments
SupervisionDid the engagement team design, perform, and review the work?Supervision record under AS 1201
DocumentationDoes the file reflect information used, output, and reasoning?Workpapers under AS 1215
Auditing client AIWhere the client uses AI in reporting, is it understood and tested?Understanding of client systems and controls

Suggested operating policy

  • AI and analytics may analyze populations, flag anomalies, extract data, and draft documentation.
  • AI output is evidence to be evaluated, never a conclusion, and never a substitute for skepticism.
  • Every exception surfaced by technology is investigated and the work is documented.
  • Data reliability and relevant IT controls are evaluated before results are relied upon.
  • Every AI tool is approved and monitored under the firm's quality-control system.
  • The engagement team supervises, reviews, and documents all AI-assisted work.
  • The auditor remains responsible for the opinion.

Thresholds and specifics should reflect the engagement's risk, the tools used, and applicable standards (GAAS, PCAOB, or ISAs).

The Changing Auditor Role

The evidence supports a shift in task composition, not the disappearance of the auditor. AI can analyze full populations, flag anomalies, extract data, and draft documentation. The auditor's work moves toward understanding data and systems, investigating exceptions, exercising calibrated skepticism, supervising the engagement, and taking responsibility for the opinion.

That shift raises the value of four abilities:

Data and technology literacy

Understanding what a tool does, where it fails, and whether its data is reliable.

Calibrated professional skepticism

Neither over-trusting nor reflexively dismissing AI output.

Investigative judgment

Turning a flagged anomaly into a supported conclusion about misstatement.

Documentation discipline

Recording the information used, the output produced, and the reasoning behind every conclusion.

The strongest future auditor is not the one who reviews samples fastest. It is the one who can supervise powerful tools, recognize when they are wrong in either direction, investigate what they surface, and stand behind the audit opinion.

Practice: Evaluate the AI Behavior

For each scenario, decide whether to Rely, Investigate, or Challenge, then reveal the best response.

Scenario 1

An analytics tool confirms that 100% of recorded sales transactions were captured and correctly totaled.

Scenario 2

An anomaly-detection model flags 300 transactions as unusual.

Scenario 3

The firm's AI system produces evidence contradicting management's complex estimate, suggesting a material adjustment.

Scenario 4

Generative AI drafts an accounting memo that reads well and cites a conclusion.

Scenario 5

A firm wants to deploy a third-party AI model whose internal logic the vendor will not explain.

Score: 0/5 correct (0 reviewed)

Key Takeaways

  • AI in auditing is real and supports full-population analysis, anomaly detection, data extraction, and documentation drafting inside firm platforms such as Deloitte Omnia.
  • "Auditing with AI" and "auditing the AI a client uses" are different activities requiring different evidence.
  • Fraud-detection models are useful but imperfect—accuracy averages roughly 78% across studies and varies widely, and strong aggregate accuracy does not guarantee catching a specific fraud.
  • Auditor reliance on AI is not always well-calibrated: research documents both automation bias (over-reliance) and algorithm aversion (under-reliance), so skepticism must be a calibration practice.
  • The standards do not create a separate AI rulebook: evidence reliability, supervision (AS 1201), documentation (AS 1215), and quality control (QC 1000) all apply to AI-assisted work.
  • The auditor remains responsible for evaluating evidence, investigating exceptions, and issuing the opinion.

The goal is not auditing without auditors. The goal is auditing in which auditors spend less time on manual sampling and more time investigating what powerful tools surface, weighing evidence with calibrated skepticism, and standing behind a defensible opinion.

Knowledge Check

Five questions on auditing with vs. the AI, research findings, algorithm aversion, standards responsibility, and responsible practice.

Question 1: What is the difference between "auditing with AI" and "auditing the AI"?

Question 2: What does the peer-reviewed research show about AI fraud-detection accuracy?

Question 3: What did Commerford et al. (2022) find about auditor reliance on AI?

Question 4: Under the auditing standards, what happens to auditor responsibility when AI is used?

Question 5: Which statement best describes responsible AI-assisted auditing?

Sources & Further Reading

Selected primary, research, and practice sources used in this module. Auditing standards and firm tools change; recheck official AICPA, PCAOB, IAASB, and vendor docs when evaluating a tool or procedure. Research accuracy figures remain tied to their study conditions.

Ready to Practice?

Take the investigation and calibration habits from this lesson into the AI Practice Arena—judgment first, automation second.

Open AI Practice Arena

What's Next?

Next in the practice catalog is AI in Financial Analysis (coming soon)—or return to the AI hub to revisit Foundations and related practice topics.

Related Concepts

Up Next

AI in Financial Analysis