Skip to main content
Concept #127

AI in Tax Preparation

Real uses, risks, rules, and examples—not marketing about a machine that files your return for you.

Educational content — not tax advice or a substitute for professional judgment. Tax law and product capabilities change; verify current IRS and vendor sources before relying on a specific capability or rule. Research accuracy figures on this page are tied to the cited study, model, and task conditions—not universal performance claims.

Why This Matters

AI is now embedded in mainstream tax preparation—but the way it actually works is narrower and more supervised than the marketing suggests. In real products, AI reads and extracts data from tax documents, auto-fills return fields, runs accuracy and completeness checks, answers plain-language questions, and helps route filers to human experts. TurboTax's Intuit Assist and the IRS's own chatbots and Interactive Tax Assistant are concrete, documented examples of these uses.

The central lesson mirrors the bookkeeping module: AI can accelerate tax preparation, but it does not transfer legal responsibility for the return away from the preparer or taxpayer. For paid preparers this is not merely good practice—it is federal law. The IRS explicitly warns that AI chatbots can fabricate information ("hallucinate") and that practitioners cannot accept AI output without independently validating it against actual tax law, court cases, and citations. Peer-reviewed and professional research confirms why this caution is warranted: general-purpose AI models make frequent, confident, and sometimes hard-to-detect errors on real tax questions.

Learning objective

By the end of this module, learners should be able to describe how AI is actually used in tax preparation, distinguish reliable automation (data extraction, calculation, e-file checks) from unreliable use (unverified legal conclusions), evaluate an AI-generated tax answer, and apply the due-diligence and data-security obligations that govern AI use in a tax practice.

What "AI in Tax Prep" Actually Means

AI in tax preparation is the use of machine-learning, document-understanding, and generative-AI capabilities to assist with preparing and filing tax returns. As with bookkeeping, it usually appears inside tax software or IRS tools rather than as a standalone "robot preparer." Documented uses fall into several distinct categories, and they carry very different reliability profiles.

Document understanding and data extraction

Reading W-2s, 1099s, and other forms and populating return fields. Intuit describes using AI-powered document understanding to read uploaded tax documents, and in 2026 announced expanded auto-fill for the ten most common U.S. tax forms using Google Cloud Document AI and Gemini models.

Accuracy and completeness checks

Scanning an in-progress return for errors, omissions, or inconsistencies. TurboTax markets real-time accuracy checks and a "CompleteCheck" scan.

Deduction and credit identification

Analyzing entered data to surface potentially applicable deductions or credits for the user to consider.

Conversational guidance

Answering plain-language tax questions. TurboTax positions Intuit Assist as an assistant that explains outcomes; the IRS deploys chatbots using natural language processing and a rules-based Interactive Tax Assistant for defined topics.

Expert routing and workflow

Matching filers to human experts and tracking progress. Intuit describes AI matching customers to its network of human experts, including CPAs, EAs, and tax attorneys, who sign returns in its full-service product.

Deterministic tax logic vs. generative prediction

The most important distinction on this page is between deterministic tax logic and generative prediction. Established tax software encodes the tax rules and formulas as programmed calculations; when the inputs are correct, the arithmetic follows the code. A generative-AI answer to a tax question, by contrast, is a probabilistic prediction of text that can be wrong even when it sounds authoritative. Conflating the two is the single most dangerous mistake a learner can make.

Where AI Enters the Return

AI changes how information flows into a return, but it does not remove any step of preparer or taxpayer responsibility.

Return stageAI-assisted activityHuman responsibility
IntakeRead and extract data from uploaded W-2s, 1099s, and other formsConfirm the document is complete, legible, and matches the taxpayer's actual situation
Data entryAuto-fill return fields from extracted dataVerify each field against the source document
Issue spottingSurface potential deductions, credits, or missing itemsDetermine whether the taxpayer actually qualifies under the law and facts
Q&AAnswer plain-language tax questionsValidate any legal conclusion against authoritative sources before relying on it
ReviewRun accuracy and completeness checks before e-filePerform the final substantive review and confirm direct-deposit and identity details
FilingRoute to a human expert where neededFor paid returns, meet due-diligence and signature obligations

The right mental model is source documents → extraction → verification → calculated return → review → filing. AI can shorten the extraction and review effort, but it cannot substitute for the verification and legal judgment that connect the return to the taxpayer's true facts.

Four Real Use Cases

1

Document extraction and auto-fill

The most reliable and mature AI use in tax prep is turning a photographed or imported document into structured return data. Intuit publicly states that AI reads uploaded documents to build a personalized checklist and auto-fill returns, and in September 2026 announced expanded extraction for complex 1099 forms (such as 1099-B, 1099-COMP, and 1099-OID) and Form 1040 schedules. This is valuable because it reduces manual transcription of many fields per brokerage document. But extraction is not verification. Reading "$4,812.55" from a 1099 correctly does not confirm that the document belongs to the taxpayer, that all documents were provided, or that the income is characterized correctly. The IRS's recordkeeping and due-diligence rules still require the preparer to evaluate whether the information is complete and consistent.

Preparer's checkpoint: Compare every auto-filled field to the source document, and confirm that the set of documents is complete for the taxpayer's situation.

2

Accuracy and completeness checks

Tax software has long used programmed logic to flag missing fields, internal inconsistencies, and common errors before e-file, and vendors now market these as AI-assisted checks. These checks are genuinely useful for catching transposed numbers, blank required fields, and e-file rejections. However, a "100% accuracy guaranteed" marketing claim refers to the vendor's calculation guarantee, not a guarantee that the return reflects the correct legal treatment of the taxpayer's facts. A return can pass every automated check and still be wrong if the underlying inputs or positions are wrong. The FTC has acted against unsubstantiated AI accuracy claims in another product category, underscoring that accuracy statements require substantiation and careful reading.

Preparer's checkpoint: Treat automated checks as a floor, not a ceiling. They confirm the return is internally consistent, not that the positions are legally correct.

3

Conversational tax guidance

Both commercial software and the IRS deploy conversational tools. TurboTax's Intuit Assist answers questions and explains outcomes; the IRS uses chatbots with natural language processing and a separate rules-based Interactive Tax Assistant that walks users through defined topics. The IRS chatbots are designed to route users to authoritative information on IRS.gov rather than to generate open-ended legal conclusions. This is the highest-risk category when the tool is a general-purpose generative model. The IRS's own Circular 230 guidance states plainly that AI chatbots are prone to fabricating facts, that there is often a lack of transparency about where an AI answer comes from, and that practitioners cannot accept AI responses—they must validate results against actual court cases, citations, and tax law.

Preparer's checkpoint: Never file a position based on an AI explanation alone. Trace every legal conclusion to the Internal Revenue Code, regulations, IRS guidance, or case law.

4

Expert matching and "done-for-you" workflows

Intuit describes an AI-driven platform that matches filers to human experts and auto-fills the majority of a return, with a human expert signing and filing in the full-service product. This design is significant precisely because a human professional remains the signing preparer and assumes responsibility for the return. That structure is not an accident. It reflects the legal reality that a paid preparer's obligations cannot be delegated to software.

Preparer's checkpoint: Automation may draft the return, but the signing preparer owns its accuracy, due diligence, and defensibility.

What the Research Actually Shows

The user-facing marketing and the empirical evidence tell two different stories. Both are true, because they describe different tasks.

General-purpose models are unreliable for tax answers

A study posted on SSRN entered common tax questions into ChatGPT (models 3.5 and 4) during the 2023 and 2024 tax seasons and found the overall share of correct responses ranged from 39% to 47%. The authors concluded the chatbot is "generally not a reliable source of tax guidance for uninformed taxpayers," and that accuracy was worse for common questions, complex answers, fact-pattern evaluation, and post-cutoff tax changes.

In Tax Notes, researchers tested GPT-4 on the SARA dataset of statutory tax cases. GPT-4 answered 186 of 276 true/false cases correctly (67%), and calculated the exact tax liability correctly only about one-third of the time, miscalculating by more than 10% in nearly a quarter of cases. Notably, the errors were not mathematical—they involved misreading the statutes. That finding is important: the model failed at legal interpretation, not arithmetic.

Accuracy depends heavily on task, domain, and retrieval

A Stanford Law white paper testing LLMs on multiple-choice tax law exams found that GPT-4—especially when combined with chain-of-thought and few-shot prompting and supplied with the correct "gold truth" legal text—could reach high accuracy, but "not yet at expert tax lawyer levels." Studies of Austrian and EU VAT law reported accuracy ranging widely with configuration; even under an idealized "perfect retrieval" setup, accuracy for determining where tax was owed reached only about 73%, and the authors concluded that a 20%–30% error rate is "unacceptable" for real practice and that an LLM-only assistant is "insufficient for practice."

Study-bound figures: Reported LLM accuracy in legal decision-making has ranged from roughly 19% to 98% depending on the domain, task type, and whether the model was purpose-built and supplied with authoritative text. A single headline accuracy number is meaningless without its task and conditions.

Hallucinated citations are a documented, sanctioned problem

The risk is not hypothetical. In the 2023 case Mata v. Avianca, attorneys relied on ChatGPT-generated research and submitted fictitious case law, resulting in judicial sanctions. A comparative analysis in The International Tax Journal reports that aggregated datasets recorded roughly 800 documented cases of AI-related citation errors across at least 25 jurisdictions by late 2025, with a sharp increase in 2025. A separate experimental study of tax-law reasoning found a strong hallucination effect affecting about 50% of analyzed court-ruling cases.

Reconciling research with product design

These findings do not mean AI is useless in tax prep—they explain why well-designed products constrain it. Established tax software performs the tax calculations through programmed logic, uses AI mainly for document extraction and workflow, and keeps a human expert as the signing preparer for complex returns. The empirical weakness of generative models on open-ended legal reasoning is precisely why the reliable uses are narrow and supervised, and why the IRS insists on human validation.

The Legal Backbone: Due Diligence, Circular 230, and Data Security

For paid preparers, AI use is governed by existing law. Nothing about AI relaxes these obligations.

Preparer due diligence

Treasury Regulation section 1.6695-2 sets out four due-diligence requirements for paid preparers claiming the EITC, CTC/ACTC/ODC, AOTC, or head-of-household status: completing and submitting Form 8867, computing credits with worksheets, not knowing (or having reason to know) that information is incorrect, and keeping specified records. The IRS is explicit that tax preparation software "helps, but it cannot replace your professional judgment or responsibility under the law." The preparer must still ask reasonable questions, evaluate whether information makes sense, and follow up on anything inconsistent, incomplete, or suspicious—and decline to prepare the return if not satisfied.

Preparers must also make additional reasonable inquiries when a knowledgeable preparer would conclude that information appears incorrect, inconsistent, or incomplete, and must document those inquiries and the client's answers at the time of the interview. These records must be kept for three years.

Circular 230 and AI

The IRS's Office of Professional Responsibility has directly addressed AI. Its Circular 230 guidance lists the standards implicated by AI use—including section 10.22 (due diligence), section 10.34 (standards on returns and documents), section 10.35 (competence), and section 10.37 (written advice)—and states that due diligence is required in assessing the reliability of AI-created results. It warns that AI chatbots fabricate facts, that there is often no transparency about the source of an AI response, and that practitioners "can't accept the responses from AI tools" but must validate them against actual court cases, citations, and tax law.

Data security is mandatory

Tax data is highly sensitive, and protecting it is a legal requirement, not an option. Federal law—under the Gramm-Leach-Bliley Act and the FTC Safeguards Rule—requires tax return preparers to create and maintain a written information security plan to protect client data, regardless of firm size. IRS Publication 4557 provides the framework, including implementing multi-factor authentication, encrypting sensitive files and emails, limiting data access to those who need it, maintaining audit logs of who did what and when, and vetting service providers so that contracts require them to maintain appropriate safeguards.

This last point is critical for AI: before entering taxpayer data into any AI tool, a preparer must know whether the vendor is a service provider under the security plan, what happens to the data, whether it is used to train shared models, and how it is protected. Unauthorized disclosure of taxpayer information can trigger civil and criminal penalties under IRC sections 6713 and 7216, which the Circular 230 guidance also flags in connection with AI use.

Worked Example: ABC Coffee Shop's Owner Files a Return

The following example is hypothetical and designed to teach verification logic, not to depict a specific software screen. The owner of ABC Coffee Shop uses AI-assisted tax software to prepare a personal return that includes Schedule C business income. Four situations arise:

SituationAI behaviorCorrect response
Uploads a 1099-NECSoftware extracts payer, amount, and taxpayer ID and auto-fills the formCompare every extracted field to the paper 1099; confirm no 1099s are missing
Asks the assistant "Can I deduct my whole car?"Assistant gives a general explanation of vehicle expense rulesDo not act on the explanation alone; verify the actual rules and the taxpayer's business-use facts and records
Software suggests a home-office deductionAI surfaces the deduction as potentially applicableConfirm the space is used regularly and exclusively for business and that the taxpayer qualifies before claiming it
Return passes the accuracy checkAutomated scan reports no errorsPerform a substantive review; the scan confirms consistency, not that positions are legally correct

Why a passing check is not a correct return

The accuracy scan can confirm the arithmetic is internally consistent and that required fields are populated. It cannot confirm that the vehicle deduction reflects actual business use, that the home office meets the regular-and-exclusive-use test, or that all income was reported. Those are factual and legal determinations that require the preparer's judgment and, for a paid preparer, documented due diligence.

A safer preparation sequence

1

Gather and verify documents

Confirm the document set is complete and belongs to the taxpayer.

2

Verify extraction

Check every auto-filled field against the source.

3

Interview and inquire

Ask about anything inconsistent, incomplete, or unusual, and document it.

4

Validate legal positions

Confirm each deduction, credit, and characterization against authoritative sources—never against an AI explanation alone.

5

Review substantively

Examine the whole return, not just the automated flags.

6

Secure the data and file

Confirm direct-deposit and identity details, protect the records, and meet signature and Form 8867 obligations where applicable.

Risks Preparers Must Control

Hallucinated or incorrect tax conclusions

A generative model can produce a confident, well-written answer that misstates the law or cites nonexistent authority. Because the errors are often interpretive rather than mathematical, they can be hard to spot.

Control: Prohibit filing any position based solely on AI output; require validation against the Code, regulations, IRS guidance, or case law, consistent with Circular 230.

Overreliance and automation bias

A polished interface or a "100% accuracy" label can discourage substantive review. The IRS is explicit that software cannot replace professional judgment.

Control: Treat automated checks as one input; require independent review of positions, deductions, and credits.

Stale or out-of-scope knowledge

Research shows models perform worse on tax changes occurring after their knowledge cutoff and on questions requiring current-year specifics. Tax law changes annually.

Control: Confirm that any AI-provided rule reflects the correct tax year, and rely on current authoritative sources for law that changed recently.

Data security and unauthorized disclosure

Entering taxpayer data into an unvetted AI tool can violate the firm's information security plan and potentially IRC sections 6713 and 7216.

Control: Vet any AI vendor as a service provider under Publication 4557; use encryption, MFA, access limits, and audit logs; and confirm data-handling and training-use terms before entering client data.

Weak documentation and audit trail

Due-diligence records must show what information was used, how it was obtained, and what inquiries were made. AI-generated content does not satisfy this on its own.

Control: Preserve Form 8867, worksheets, source documents relied upon, and contemporaneous notes of inquiries and answers for three years.

A Control Framework for AI in a Tax Practice

A small firm does not need an elaborate governance committee, but it does need explicit decisions about scope, verification, security, and records. The structure below mirrors NIST's approach to AI risk—establish accountability, understand context, measure performance, and manage risk—applied to tax.

Control areaPractical questionMinimum evidence
ScopeWhich tax tasks may AI perform (extraction, checks, drafting) versus not (final legal conclusions)?Approved-use policy
VerificationWho confirms extracted data and legal positions?Review sign-off
AuthorityCan the tool draft, but never file, without human review?Workflow and permission rules
Data securityIs the AI vendor vetted under the firm's Publication 4557 plan?Written security plan and vendor terms
Due diligenceAre Form 8867, worksheets, and inquiry notes maintained?Retained records for three years
ValidationIs every AI legal conclusion traced to authority?Citation to Code, regs, guidance, or cases
CurrencyDoes the guidance reflect the correct tax year?Confirmation against current sources
MonitoringHow are errors detected and corrected over time?Error log and periodic review

Suggested operating policy

  • AI may extract data, populate fields, run checks, and draft explanations.
  • AI may not be the sole basis for any filed tax position or legal conclusion.
  • Every auto-filled field is verified against source documents.
  • Every deduction, credit, and characterization is validated against authoritative law.
  • No taxpayer data is entered into an AI tool that is not vetted under the firm's information security plan.
  • All due-diligence records are completed and retained for three years.
  • The signing preparer remains responsible for the return.

Thresholds and specifics should reflect the firm's clients, risk, and the requirements of the credits and positions involved.

The Changing Tax-Preparer Role

The evidence supports a shift in task composition, not the disappearance of the preparer. AI can read documents, populate fields, flag inconsistencies, and draft explanations. The preparer's work moves toward verifying inputs, interviewing clients and documenting inquiries, validating legal positions, protecting data, and taking responsibility for the return.

That shift raises the value of four abilities:

Tax-law knowledge

Understanding the Code, regulations, and current-year rules well enough to catch an AI error.

Professional skepticism

Treating a confident AI answer as a claim to be verified, not a conclusion to be filed.

Due-diligence discipline

Asking the right questions, evaluating answers, and documenting them.

Data stewardship

Protecting taxpayer information as both an ethical duty and a legal requirement.

The strongest future preparer is not the one who types return data fastest. It is the one who can supervise the software, recognize when it is wrong, protect the client's data, and stand behind the return.

Practice: Evaluate the AI Behavior

For each scenario, decide whether to Rely, Verify, or Reject, then reveal the best response.

Scenario 1

The software extracts a W-2 and auto-fills wages and withholding. The figures match the paper W-2 exactly.

Scenario 2

A chatbot answers, "Yes, you can deduct all of your commuting miles as a business expense."

Scenario 3

The assistant cites a specific court case supporting an aggressive deduction, but the preparer cannot locate the case.

Scenario 4

The return passes the automated accuracy check with no flags, but the client's reported business expenses look inconsistent with their income.

Scenario 5

A preparer wants to paste a client's full tax documents into a public AI chatbot to speed up drafting.

Score: 0/5 correct (0 reviewed)

Key Takeaways

  • AI in tax preparation is real and already supports document extraction, auto-fill, accuracy checks, conversational guidance, and expert routing inside products like TurboTax and IRS tools.
  • Deterministic tax calculations are not the same as generative predictions; the reliable uses are narrow and supervised for good reason.
  • Research shows general-purpose models are frequently wrong on real tax questions—one study found 39%–47% accuracy, and another found exact liability correct only about a third of the time—with errors that are interpretive and hard to detect.
  • Hallucinated citations are a documented, sanctioned risk, not a hypothetical one.
  • For paid preparers, AI does not relax due-diligence, Circular 230, or data-security obligations; the IRS requires human validation of AI output and a written information security plan.
  • The safest operating model combines narrow automation, verification of inputs, validation of positions against authority, strong data security, documented due diligence, and human responsibility for the return.

The goal is not tax preparation without professionals. The goal is tax preparation in which professionals spend less time transcribing documents and more time verifying facts, validating positions against the law, protecting client data, and standing behind an accurate return.

Knowledge Check

Five questions on reliable AI uses, research findings, Circular 230 validation, data security, and responsible practice.

Question 1: Which tax-prep task is generally the most reliable AI use?

Question 2: What did research find about general-purpose models answering tax questions?

Question 3: Under IRS guidance, what must a practitioner do with an AI-generated tax answer?

Question 4: What is legally required before a paid preparer enters client data into an AI tool?

Question 5: Which statement best describes responsible AI-assisted tax preparation?

Sources & Further Reading

Selected primary, research, and product sources used in this module. Tax law and product pages change; recheck official IRS and vendor docs when evaluating a tool or position.

Ready to Practice?

Take the verification habits from this lesson into the AI Practice Arena—judgment first, automation second.

Open AI Practice Arena

What's Next?

Next in the practice catalog is AI in Audit (coming soon)—or return to the AI hub to revisit Foundations and related practice topics.

Related Concepts

Up Next

AI in Audit