Evaluating AI Tools as a CPA
How to vet vendors, data privacy, and client risk before a cool demo becomes a firm-wide problem.
Educational content — not legal, cybersecurity, privacy, tax, or professional-standards advice. Soft-label AICPA, NIST, IRS, FTC, and CPA.com claims on this page as regulator, standards, or vendor documentation unless independent evidence is cited. Confirm requirements with qualified counsel and your firm’s policies before adopting a tool.
Product freshness: Verified September 18, 2026. AI products, contracts, model providers, privacy terms, regulations, client expectations, and professional guidance change quickly. Review the vendor record before each material new use case and reassess at least annually.
Why This Matters
A useful demo is not a risk assessment. Before a CPA firm or accounting team adopts an AI tool, it should know what the tool does, what data enters it, where that data goes, who can access it, whether it is retained or used for training, what the vendor promises contractually, how the tool is tested on firm-specific work, and what happens if the tool fails or changes.
The core idea
Cool demos are not due diligence. Specific use, specific data, specific contract, specific controls—or do not adopt.
Learning Objectives
By the end of this lesson, you should be able to:
- Separate a product demo from real vendor due diligence.
- Map the data flow of an AI tool before client or firm data is entered.
- Ask practical questions about retention, training, access, subprocessors, encryption, location, model changes, and incident response.
- Evaluate whether a proposed use case is appropriate for confidential client information.
- Test an AI tool using safe, representative data and documented success criteria.
- Connect AI-tool evaluation to engagement letters, client expectations, confidentiality, and applicable tax-data restrictions.
- Build a repeatable approval process for low-, medium-, and high-risk tools.
- Monitor a vendor after adoption rather than treating procurement as a one-time event.
The Question Is Not “Is This Tool Good?”
Ask this instead
Is this specific use of this specific tool, with this specific data, under this specific contract and control environment, acceptable for this firm and this client engagement?
Part 1
Use case
What exact problem will the tool solve—not “AI for the tax department,” but a specific workflow with defined inputs and outputs.
Part 2
Data
What information will be submitted, retrieved, retained, or generated—including prompts, uploads, logs, embeddings, and exports.
Part 3
People
Who can use it, administer it, see logs, and approve output—and who is accountable when it fails.
Part 4
Technology
Which vendor, model, hosting environment, plug-ins, connectors, and subprocessors are involved.
Part 5
Contract
What the agreement says about data use, ownership, confidentiality, security, liability, audit rights, and incident response.
Part 6
Professional obligations
Whether the use aligns with confidentiality, engagement scope, tax-data restrictions, independence where applicable, and client expectations.
Part 7
Controls
How the firm will test, supervise, document, and monitor the workflow after approval.
NIST AI Risk Management Framework materials are voluntary guidance, but they offer a useful structure: Govern, Map, Measure, and Manage. Those materials call for policies addressing risks from third-party software and data, mapping risks across third-party components, contingency planning for high-risk third-party failures, and regular monitoring of third-party AI risk controls.
Start with a Use-Case Statement
Do not evaluate “AI” as a category. Evaluate a specific proposed workflow.
Weak use-case statement
“We want a chatbot for the tax department.”
Too vague. It does not say what work, what data, who reviews, which environment, or what is prohibited.
Build an AI-Tool Inventory
Every tool under evaluation or in use should have an inventory record. The inventory is the starting point for governance, not bureaucratic paperwork.
| Field | What to record |
|---|---|
| Tool and vendor | Product name, legal vendor entity, product edition, URL |
| Business owner | Person accountable for the use case |
| Technical owner | IT, security, data, or system administrator |
| Proposed use case | Specific task; avoid broad descriptions such as “assist with accounting” |
| Users | Roles, teams, client-service staff, external users if any |
| Data classification | Public, internal, confidential, client confidential, tax-return information, personal information, financial-account data |
| Data flow | Inputs, uploads, retrieved data, outputs, logs, exports, integrations |
| Model and hosting | Vendor model, third-party model/provider if disclosed, region/environment |
| Subprocessors | Known hosting, model, OCR, support, analytics, or other subprocessors |
| Output consequence | Low, medium, or high; include whether output affects client advice, accounting records, filings, payments, reports, or audit evidence |
| Human review | Who reviews, at what stage, and what they must verify |
| Contract status | DPA, confidentiality terms, service terms, SLA, security exhibits, audit rights |
| Testing record | Test data, criteria, results, limitations, approval date |
| Change-monitoring plan | Trigger events, review frequency, responsible owner |
| Exit plan | Data export/deletion, replacement process, business-continuity plan |
NIST materials recommend documenting the scope and context of an AI application—including risks across third-party components—and reviewing vendor audit reports, testing results, contracts, release schedules, and change-management plans.
Map the Data Before the Demo
The first serious question is not “Does it have SOC 2?” It is “Where does our data actually go?”
Client file / firm system
↓
User prompt, upload, connector, or API
↓
AI application / vendor environment
↓
Model provider, OCR provider, cloud host, analytics, support, or other subprocessors
↓
Output, logs, chat history, exports, retention, backups
↓
Firm workpaper, client deliverable, or accounting systemCPA.com’s AI Solution Due Diligence Guide recommends asking vendors for a clear data-flow explanation that identifies what data is accessed, where it is processed and stored, whether it is sent to an external large language model, and what is redacted or anonymized before it leaves the system.
Data-flow questions
Checked 0/11 data-flow questions
Client Confidentiality Comes First
AICPA Code of Professional Conduct materials address use of third-party service providers. Before disclosing confidential client information to a third-party service provider, a member should either enter into a contractual agreement requiring the provider to maintain confidentiality and provide reasonable assurance of appropriate procedures to prevent unauthorized release, or obtain the client’s specific consent. Those materials also say the member should inform the client, preferably in writing, that a third-party service provider may be used; if the client objects, the member should not use the provider or should decline the engagement.
An AI vendor may be a third-party service provider when it receives or can access confidential client information. The label “AI,” “assistant,” “copilot,” or “productivity tool” does not change the underlying confidentiality question.
Practical rule: Never assume that a free account, consumer account, personal browser extension, or employee-purchased tool is appropriate for client data just because the user can log in. Approval must depend on the actual account tier, contract, settings, integrations, data flow, and engagement context.
AICPA materials note that firms should expect clients to ask how information is stored, whether it is used to train a model, who can access it, whether it is shared outside the firm, how it is retained, and whether de-identification procedures exist.
Tax Data Is a Special Case
Tax return information is not simply “confidential data.” It can be subject to specific legal restrictions. According to IRS materials, Internal Revenue Code section 7216 generally prohibits tax return preparers from disclosing or using tax return information for non-return-preparation purposes without taxpayer consent, subject to statutory and regulatory exceptions. IRS materials further explain that consent must be knowing and voluntary and that tax preparers should clearly inform taxpayers with whom information will be shared and for what purpose.
Practical implication — determine with legal and compliance advice as needed:
- Whether the proposed use is permitted without taxpayer consent.
- Whether the vendor is functioning as a permitted service provider under the applicable rules.
- Whether consent is required for the specific use or disclosure.
- Whether the firm’s engagement letter, privacy notice, consent language, and internal policy address the intended workflow.
- Whether the tool receives tax-return information directly, indirectly through a connector, or through output logs and support tickets.
This page is educational and does not provide legal advice. Section 7216 questions are fact-specific and should be reviewed with qualified legal, tax, and compliance professionals.
Security Due Diligence
Security evidence matters, but no single certification answers every question. A SOC 2 report, for example, may be helpful evidence about selected controls during a stated period; it does not itself prove that a particular AI workflow is safe, that data is not used for training, or that outputs are accurate.
CPA.com’s due-diligence guide recommends requesting documents such as a SOC 2 report (preferably Type 2), service-level agreement, user agreement, subprocessor list, privacy policy, data-processing addendum, data-flow architecture diagram, incident-response plan, historical incident information if available, and accuracy-benchmark summary.
| Topic | Questions to ask |
|---|---|
| Identity and access | Does the tool support SSO, MFA, role-based access, least privilege, and administrator controls? |
| Encryption | Is data encrypted in transit and at rest? What exceptions exist? |
| Tenant isolation | How is one customer’s data separated from another customer’s data? |
| Logging | What user, admin, model, and data-access events are logged? Can the firm export logs? |
| Incident response | How quickly will the vendor notify the firm? Who investigates? What support is provided? |
| Vulnerability management | How are security vulnerabilities, patches, model issues, and dependencies tracked and remediated? |
| Subprocessors | Who are they, what do they do, and how will changes be disclosed? |
| Data deletion | Can the firm delete prompts, uploads, outputs, and account data? How are backups handled? |
| Business continuity | What happens if the service is unavailable, discontinued, acquired, or changes materially? |
| Audit evidence | Can the vendor provide relevant independent assurance reports, policies, or contract commitments? |
For entities covered by the FTC Safeguards Rule, FTC materials describe a written information-security program with administrative, technical, and physical safeguards appropriate to the organization and the sensitivity of customer information, including oversight of service providers. Applicability depends on the firm and facts—obtain appropriate legal or compliance advice rather than assuming the Rule applies or does not apply.
Training, Retention, and Model Improvement
“Your data is private” is not a complete answer. Ask the vendor to separate these questions:
- Is data used to respond to the current request?
- Is data retained in chat history, logs, diagnostics, abuse monitoring, or backups?
- Is data accessible to vendor personnel or subcontractors for support, safety, or service operations?
- Is data used to train, fine-tune, evaluate, or improve a general or customer-specific model?
- Can the firm opt out, and is the opt-out contractual or merely a product setting?
- Does the answer change by product tier, geography, model, connector, or feature?
Never accept these as final answers
Those statements may be true, false, incomplete, conditional, or different across products. Ask which data, which feature, which plan, which region, which model, which subprocessor, and which written commitment.
Accuracy Is Not the Same as Usefulness
A model can create impressive text and still be unsuitable for an accounting use case. Define what “good enough” means before testing.
| Risk level | Example use | Adoption threshold |
|---|---|---|
| Low | Draft internal meeting agenda from nonconfidential notes | Basic quality review and approved environment |
| Moderate | Summarize an approved management-report narrative | Representative testing, human review, output controls |
| High | Extract client source documents, draft tax analysis, classify transactions, or prepare workpaper content | Formal data review, contract review, representative testing, documented approval, enhanced monitoring |
| Very high | Generate journal entries, affect a filing, recommend a tax position, influence payment, support an audit conclusion, or make a professional conclusion | AI output cannot be the sole basis; use case requires qualified human responsibility, formal controls, and potentially legal/compliance review |
NIST materials warn that generative AI can confabulate—produce confidently stated but false content—and recommend evaluating capability claims with empirically validated methods instead of extrapolating from narrow or anecdotal demonstrations.
Test Before Adoption
A demo uses curated facts. A pilot should use representative work. Prefer firm-created fictional data, then de-identified approved historical data, then controlled client samples only after contract, consent, and security requirements allow it. Do not test using data that production itself would prohibit.
Include difficult cases
Normal cases; blank and incomplete records; duplicates; ambiguous names; multiple entities or currencies; exceptions; conflicting source documents; outdated policy references; poor scans/tables/handwriting; and prompts designed to test whether the tool invents facts, citations, or explanations.
Test output categories
| Result | Meaning | Next action |
|---|---|---|
| Correct and supported | Output matches approved ground truth and cites/uses permitted inputs | Consider approval if controls also pass |
| Correct but not explainable | Output appears right but source or method cannot be reviewed | Restrict use; do not use for high-consequence work |
| Partially correct | Some useful elements but omissions, ambiguity, or inconsistent reasoning | Redesign prompt/workflow; add human review or reject use case |
| Incorrect but obvious | Error is easy for a reviewer to spot | May be acceptable only for low-risk drafting with controls |
| Incorrect and plausible | Error could escape ordinary review | High risk; reject or redesign workflow |
| Unsafe | Leaks data, violates policy, invents authorities, changes records, or produces prohibited output | Stop pilot; escalate |
Evaluate the Human Workflow Too
A tool is not safe simply because it is accurate in a test. The entire work process matters.
- Will staff understand when the tool is making a suggestion rather than stating a verified fact?
- Can users see the source documents or data behind an output?
- Does the interface make it easy to accept a recommendation without reviewing it?
- Who reviews output, and do they have enough time and expertise to challenge it?
- What happens when the tool is uncertain, wrong, unavailable, or updated?
- Can staff override or correct it? Is the correction captured for future review?
- Does the workflow preserve a workpaper-quality audit trail?
- Does the tool encourage users to upload more client data than is needed?
NIST AI RMF materials call for processes for human oversight to be defined, assessed, and documented, and for deployed systems to be monitored for validity and reliability.
Vendor Questions That Matter
Use this as a live due-diligence script—expand each category.
- What precise problem does the tool solve for accounting or CPA firms?
- Which features are generally available, in preview, beta, or dependent on a paid tier?
- Which features use generative AI, machine learning, OCR, rules, external web search, agents, or third-party foundation models?
- What are the known limitations, unsupported tasks, and common failure modes?
- What changes have been made to the model, system, pricing, or data handling in the last 12 months?
- How are future material changes communicated?
Engagement Letters and Client Expectations
Engagement letters define scope, deliverables, timelines, and responsibilities. They are not a one-line solution to AI risk. Whether AI use needs disclosure, client notice, consent, or revised terms depends on the facts, governing law, professional requirements, service model, data involved, and client agreement.
AICPA materials say no single federal law or professional standard universally mandates disclosure whenever a CPA uses generative AI, but emphasize the patchwork of legal, privacy, security, ethical, and client-expectation considerations—and recommend consulting legal counsel when drafting AI disclosure or consent language.
- Does the engagement letter permit use of third-party service providers?
- Does the work require sharing confidential client information with the AI vendor or its subprocessors?
- Has the client been informed in the manner required by applicable professional obligations and contract terms?
- Does the client have a policy prohibiting certain AI tools, countries, cloud services, or data uses?
- Is the firm using AI only for internal administrative work, or is it part of the professional service delivered to the client?
- Does the tool change the scope, timeline, staffing, review process, or responsibility described in the engagement letter?
- Are consent and disclosure questions different because tax-return information is involved?
Contract Review: Find the Real Promises
Do not evaluate privacy or security based solely on a sales presentation or help-center article. Identify the contractual documents that govern the firm’s actual purchase:
Contract red flags
- The vendor may use prompts, uploads, outputs, or derived data for broad product improvement without a clear opt-out.
- Data ownership or output rights are vague.
- The vendor can change data terms unilaterally without meaningful notice.
- The vendor disclaims all responsibility for security incidents while receiving sensitive data.
- Breach-notification timing is unclear.
- There is no commitment to delete or return data after termination.
- Subprocessors can change without notice or are not identified.
- The firm cannot obtain enough evidence to meet its own client or regulatory obligations.
- The contract prohibits the firm from testing, auditing, or assessing the tool.
- The use case depends on an experimental feature with no stable support, retention, or security commitments.
Make Approval Proportional to Risk
Not every tool needs a months-long procurement process. But high-consequence tools deserve more than a click-through approval.
Tier 3: High
Example: Work with confidential client data, accounting exports, workpapers, or non-tax personal information
Required approval approach: Security/privacy review, contract review, representative test, client/engagement assessment, formal approval, monitoring plan
A 30-Day Pilot Plan
Define → test safely → design controls → decide.
Week 1 — Define
- Write the specific use case.
- Identify owner, users, output consequence, and prohibited data.
- Map the initial data flow.
- Review the engagement and confidentiality implications.
- Collect vendor documentation and contracts.
A pilot that does not define what would cause rejection is not a real test—it is a slow-motion rollout.
Monitoring After Approval
Vendor risk changes when the vendor changes the product, model, data handling, feature set, integration, pricing, subprocessor, or contractual terms. Monitor the use case, not only the vendor brand.
Reassessment triggers
- New AI model or major product release.
- Change in training, retention, data-residency, or subprocessor terms.
- New connector, agent, browser extension, web-search feature, or API integration.
- New use case or new data classification.
- Security incident, outage, material error, or client complaint.
- New regulatory, legal, contractual, or professional requirement.
- Acquisition, financial distress, or major organizational change at the vendor.
- Annual review even if no change is announced.
NIST materials call for regular monitoring of third-party AI resources and documentation of applied risk controls, as well as contingency processes for high-risk third-party system failures or incidents.
The Exit Plan
Ask before adoption: What happens if we need to stop using this tool tomorrow?
- How to export firm data, prompts, outputs, workpapers, and audit logs.
- How to revoke users, API keys, connectors, tokens, and service accounts.
- How to verify deletion or return of data where contractually available.
- How to preserve records required for engagements, quality control, legal hold, retention, or regulatory purposes.
- How to complete work if the tool is unavailable during busy season or close.
- How to replace AI-generated workflow components with manual or alternative controlled processes.
A vendor dependency is a business-continuity issue, not merely an IT issue.
ABC Coffee Shop: A Practical Evaluation
Walk an invoice AP coding-queue evaluation from workflow to claim discipline.
Proposed AP coding queue
ABC Coffee Shop’s outside CPA firm wants an AI document tool to read vendor invoices and create a proposed accounts-payable coding queue.
- Staff upload invoices from the client’s approved document repository.
- The tool extracts vendor name, invoice number, dates, line items, tax, and total.
- The tool proposes a vendor match and expense-account code.
- An AP specialist reviews every proposed record before entry into the accounting system.
- The tool cannot create vendors, change bank details, approve payments, post bills, or release payments.
Common Mistakes
| Mistake | Why it is risky | Better practice |
|---|---|---|
| Approving a tool after a polished demo | Demos rarely show real data, edge cases, contract limits, or failure modes | Pilot against documented, representative criteria |
| Treating a privacy page as the contract | Marketing language may not govern the purchased service | Review the order form, DPA, service terms, and AI-specific terms |
| Asking only “Do you train on our data?” | Retention, support access, logs, providers, and feature-specific use may still matter | Map all data uses and get written answers |
| Letting staff use personal or free accounts | Account tier, identity, and settings may lack required protections | Use only approved firm-managed accounts and configurations |
| Treating a SOC report as a full approval | It may not cover the feature, model, workflow, data use, or output risk | Combine assurance evidence with data-flow, contract, and pilot review |
| Testing only simple examples | Real accounting work includes ambiguity, exceptions, and bad documents | Include representative difficult cases and define ground truth |
| Accepting a “confidence score” as proof | A vendor score is not independent evidence or an accounting conclusion | Require source review, human approval, and measurable tests |
| Adding generic AI language to an engagement letter | Broad boilerplate may not match the actual tool, data flow, or consent obligation | Use accurate, fact-specific language with counsel where appropriate |
| Ignoring product updates | The approved workflow can change when the vendor changes | Define reassessment triggers and monitor changes |
| Having no exit plan | Outage, breach, termination, or vendor change can disrupt client work | Plan export, deletion, replacement, and continuity before adoption |
Vendor Evaluation Worksheet
Interactive A–E checklist you can use as a firm approval template.
A. Use case
B. Data map
C. Professional and contract review
D. Test record
E. Approval and monitoring
Practice: Rely, Verify, or Reject
Choose Rely, Verify, or Reject for each scenario, then reveal the model answer.
Scenario 1: A vendor says, “We are SOC 2 compliant, so you can upload any client file.”
Scenario 2: A tool is approved for drafting internal policy summaries from public accounting standards. A staff member wants to paste in a client’s full tax return to ask a question.
Scenario 3: The vendor says, “We do not train on enterprise customer data,” but the contract allows retention of prompts and outputs for 30 days to provide support and detect abuse.
Scenario 4: An AI invoice tool performs well on twenty clean invoices from one vendor but fails on handwritten credit memos, duplicate invoice numbers, and multi-currency invoices.
Scenario 5: The engagement partner says, “The client has never objected to our cloud software, so no AI review is needed.”
Knowledge Check
Five questions on use-case definition, data uses beyond training, meaningful pilots, engagement letters, and ongoing monitoring.
Question 1: What is the first question to ask before adopting an AI tool?
Question 2: Why is “we do not train on your data” not enough?
Question 3: What makes a pilot meaningful?
Question 4: Can a client engagement letter automatically solve AI confidentiality questions?
Question 5: When does vendor due diligence end?
Final Adoption Checklist
Key Takeaways
- A useful demo is not a risk assessment—evaluate a specific use of a specific tool with specific data under a specific contract and control environment.
- Map data flow before the demo: inputs, transfers, subprocessors, retention, access, and contractual protections.
- Client confidentiality and tax-return information rules can make consumer or free accounts inappropriate for firm work.
- “No training” is incomplete without answers on retention, support access, logs, location, deletion, and feature-specific behavior.
- Define success criteria before testing; include difficult cases; categorize incorrect and unsafe outputs.
- Make approval proportional to risk, then monitor after approval—vendor risk changes when the product, model, terms, or features change.
- Always have an exit plan for export, revocation, deletion, and busy-season continuity.
A cool demo can sell a seat. Only a documented use case, data-flow map, contract review, representative pilot, and proportional approval sell a responsible firm decision.
Sources & Further Reading
Selected NIST, AICPA, IRS, FTC, and CPA.com materials. Product and guidance pages change; recheck official sources when evaluating a tool.
- NIST — Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- NIST — Artificial Intelligence Risk Management Framework: Generative AI Profile (NIST AI 600-1)
- AICPA Code of Professional Conduct (ET) — confidentiality & third-party service providers
- AICPA & CIMA — generative AI resources for CPAs (professional insights)
- IRS — Disclosure or use of tax return information by preparers (IRC §7216 overview)
- IRS — Section 7216 regulations and consent guidance (tax pros)
- FTC — Safeguards Rule (financial institutions & customer information)
- CPA.com — AI Solution Due Diligence Guide
What's Next?
Continue with career and ethics lessons, or deepen practical AI workflows already live.