WorkpaperIQ · Client-facing · v1.0 · 2026-09-03

AI: Client Questions & Answers

This is reference material for conversations about AI in the WorkpaperIQ service. Each entry leads with the direct answer in bold, followed by the context you are likely to need for the follow-up question. Questions that clients ask in several different ways have been consolidated, so the same question always gets the same answer.

The last section is different from the rest: it covers the questions your own audit clients are likely to ask you once they know an AI tool touched their engagement, and what we can put in your hands to answer them.

01How WorkpaperIQ uses AI 02Confidentiality, security & governance 03Your control & transparency 04Human oversight & quality 05Pricing & efficiency 06Talking to your own audit clients

01 · How WorkpaperIQ uses AI

What the technology does, which models are behind it, and where ordinary software ends and a language model begins.

What exactly does the AI do on my engagement?

It reads your completed workpapers and writes review comments for a human to judge. The service reads the package you submit, classifies each file by its role and relevance, works through the audit area's review procedures step by step, and reports findings in four classes — Critical Finding, Recommended Correction, Documentation Enhancement, and Required Audit Team Decision — each anchored to a file, a tab, and where determinable a cell. It does not perform audit procedures, it does not write in your workpapers, and it produces no opinion or conclusion.

Which AI models are you using?

Two production model providers, through inference APIs only. An enterprise LLM inference endpoint runs the primary review engine and the second-pass review (provider and contracting entity named to your vendor-risk reviewer on request); the Anthropic Claude API is used for parts of the analysis, the in-product help assistant, and internal quality scoring. A third provider (DeepSeek) is used for internal development experimentation only and is barred by policy from production customer data; client analyses are technically pinned to the production engine, and a technical enforcement review covering every path is on our roadmap. Every provider is listed as a subprocessor in our vendor register, which we provide on request. Our vendor policy sets a minimum of 30 days' advance notice before a new material subprocessor processes customer data; that commitment takes effect with our first production Order Form.

Is all of it AI, or is some of it ordinary software?

A meaningful part is deterministic code, and that distinction matters for your risk assessment. Reading a spreadsheet, resolving merged cells, extracting a formula, checking a citation against the files you actually submitted, enforcing a severity downgrade, deleting a file after the run — those behave identically every time. Language models are used for classification, the review reasoning, the second-pass review, and the adversarial re-check.

Does the help assistant in the product read our workpapers?

No. The in-product assistant sees interface state only — which screen you are on, how many findings the session has, which filters are active, how much credit remains, and the names you gave the session and its files. It never receives workpaper contents or finding text, and it answers questions about using the site rather than audit questions. Each question is relayed through our server to the model provider to be answered; we do not store the conversation, with one exception: if you rate an answer with the thumbs control, we keep that question and answer to improve the assistant.

Who at our firm can run it?

Whoever you issue an access code to, within the role you choose. Codes are issued per firm and per role — preparer or reviewer. A preparer seat can be scoped so that preparer sees only the sessions they created; reviewer codes see the firm's work. Supervision is unchanged by the technology: the person who dispositions a finding is named on it.

02 · Confidentiality, security & governance

The section vendor-risk reviewers read first. Where the honest answer is narrower than the industry-standard one, it is written narrower.

Is our data used to train AI models?

No. We do not train or fine-tune any model on customer data, we build no datasets from it, and we do not use workpapers, findings, or engagement data in prompt development or benchmarking (our benchmark packages are fictional workpapers with planted deficiencies). Our production model providers are used through inference APIs under enterprise terms that commit them not to train on the data we send.

For your vendor-risk file. We hold no-training, inference-only commitments from each production provider and will make the confirmations available to your reviewer on request, alongside our subprocessor register.

Where is our data stored?

On AWS, in Seoul today; a U.S. region is a committed roadmap item. Production runs in ap-northeast-2. Migration to us-east is committed before the first paid production engagement, and the residency commitment takes effect as stated in your Order Form. If U.S. residency is a precondition for your firm, say so and it becomes a gating item on your onboarding rather than a promise in a footnote. Residency describes where we host your data; workpaper excerpts are additionally sent to our model providers' inference APIs for processing, in the regions those providers operate.

Is our data encrypted?

In transit, yes. At rest, not yet verified — and we say so. All public endpoints serve traffic over TLS. Encryption at rest on our storage volumes is unverified today; verifying and, if needed, enabling it is a dated roadmap item (2026-09-15), as is multi-factor authentication on our infrastructure accounts. These are the two questions every vendor-risk questionnaire opens with, so we would rather answer them here than have you find the gap yourself.

How long do you keep our workpapers?

The working copy is deleted right after the run; results and an organisation-scoped archived copy of the workpapers persist for the life of the engagement plus 60 days. The archive exists so that a re-run or a fix-verification does not require you to upload the same binder again; it is scoped to your firm and purged on request, along with the session's results and outputs. This is our starter retention schedule — a complete written schedule covering every store is on our roadmap.

Could another firm's data ever reach our binder, or ours theirs?

Between firms, no — enforced server-side. Within your firm, an entity gate keeps engagements apart, with stated limits. Every session, binder, and credit pool is scoped to one firm. Within a firm, an engagement is bound to the audited-entity names the service reads on the first run, shown on the engagement's scope line; a later package naming a different entity stops at the classification step for human approval before any review work is done, and you are charged only for that classification step. The gate needs a recognisable entity name to act on — unlabelled scans or schedules will not trigger it — and it applies to engagement runs, not to stand-alone sessions.

Who at your company can see our workpapers?

One engineering administrator, for support and operations. Operational logs, reports, and dashboards show masked email addresses and masked access codes rather than raw values. We would rather tell you that one person has that access than imply nobody does.

Do you hold SOC 2 or ISO 27001?

No, and we do not represent otherwise. We publish our security, privacy, vendor, incident-response, and change-management policies in full, and we answer vendor-risk questionnaires in detail. If a certification is a hard requirement for your firm, tell us early — we would rather be told at evaluation than discovered at contracting.

What stops someone hiding an instruction inside a workpaper to manipulate the review?

Untrusted-data framing, conclusion masking, and citation grounding. Free-text notes attached to a run are presented to the model as untrusted data that cannot change the review rules. The workpaper's own conclusions are masked during review, so the model re-derives its analysis rather than agreeing with the preparer. A finding whose cited evidence cannot be matched to a file you actually submitted is downgraded and routed for human decision instead of being presented as a confident call.

How is a change to the AI controlled?

Benchmark regression before production, every time. Any change to a model, a prompt, or the review logic is run against benchmark sets of fictional workpapers with known planted deficiencies and scored against the prior baseline. Detection of critical planted deficiencies must hold — a change that reduces it does not ship. A movement outside our tolerance band on lower-severity measures is escalated and can be released only on documented approval, with the deviation recorded in the release note. Changes reach a data-isolated development environment first, and production deployment happens only when no customer analysis is running.

03 · Your control & transparency

What your firm can require, restrict, and get back.

Can we restrict or prohibit AI use on a particular engagement?

Yes — by controlling what the service ever sees. WorkpaperIQ is an AI product end to end, so there is no "run it without the AI" setting; that would be a different product. What you control is everything upstream of it: which engagements and audit areas go through the service at all, which individual files are ticked for a given run (the tick list is shown before the run starts, and you can untick anything), which entities an engagement is bound to, who at your firm holds a code, and deletion on request. If one of your audit clients prohibits AI on its engagement, the honest implementation is that the engagement is not run through the service — and we will document that restriction at the engagement level.

If our client's terms conflict with your policy, which controls?

Your client's terms. We will not ask you to reconcile a conflict between our policy and an obligation you owe your client. Tell us the constraint and we document and honour it.

What do you tell us about what actually happened in a run?

A coverage record, per file and per tab. Every analysis reports which files and tabs were read in full, which were read only in part and to what row, which were not read and why, and which could not be opened at all. That record appears on the results page, in the exported workbook, and in a summary line at the top of the report.

Why this is stated so precisely. A finding that stands on a file the service could not read is collapsed into one "could not be read" card and routed for manual review — never presented as a missing workpaper. That distinction was one of the two root causes behind the false findings in our first real-workpaper test, and closing it is the reason the coverage record exists.

Can we get the decision trail out of the system?

Yes. Each finding carries its disposition, the name or initials of the person who made it, any note they wrote, and the timestamp. The workbook export and the review-record export both carry that trail, and they are yours to file in the engagement.

Will you tell us if something goes wrong?

Yes, under a written incident-response policy. Our Security Incident Response Plan sets out how we triage, contain, and notify. Notification commitments to your firm live in the engagement documents; the operational detail lives in the plan, which we will send on request.

04 · Human oversight & quality

What the product enforces, what it cannot do, and how good it actually is.

Is this replacing our reviewers?

No. The service reads the files you submit that are in scope for the area you picked, applies that area's review procedures to every one of them, and reports exactly what it read, what it read only in part, and what it set aside. It does not exercise professional judgment, does not assume responsibility, and does not sign anything. Every finding requires a human disposition before it counts as resolved, the disposition is recorded against a named person, and the product never writes in a workpaper.

How do you keep it from inventing findings?

Two adversarial re-checks, with different powers. Before a Critical or Recommended finding that cites locatable evidence reaches you, a separate pass tries to refute it against that evidence. If it positively contradicts the finding — at a high confidence bar, and only against evidence it actually located — the finding is withdrawn and the withdrawal is recorded in the run record, which you can ask for. If it finds the severity overstated, the finding is downgraded one level. A Critical claim that something is missing from the package is treated more conservatively: up to eight such claims per run — those citing the least evidence first — trigger a search of the whole binder, and that check can only downgrade the finding one level with the counter-evidence attached — it can never remove it; claims beyond the cap fall back to the ordinary re-check. We built it that way because "it is missing" is the claim most likely to be wrong for the wrong reason, and you should see it with the evidence rather than not at all.

How accurate is it, and how do you know?

Measured against benchmark packages with known planted deficiencies, before every release. We maintain fifteen landmine sets across audit areas plus clean sets that measure over-reporting. On the most recent release, detection of critical planted deficiencies was 208 of 209 across the landmine sets, and false findings on our clean precision package fell from 27 to 12 — it still raises about a dozen items on a package with nothing wrong in it, which is one reason every item requires your disposition. These are our own benchmarks on fictional packages, not an independent assessment — we describe them as what they are, and the methodology will be shareable under NDA (target 2026-10-15).

What is it bad at?

We publish that separately, and we would rather you read it before deployment. The Capabilities & Known Limitations document lists the material limitations. In short: it reviews what you give it — it cannot know about a workpaper that was never uploaded; very large detail files can exceed the reading budget for a single run (the coverage record tells you when that happened); scanned material depends on OCR quality; results can vary between runs on identical input and can change after a model or prompt update; and judgment-heavy matters are exactly where its output is a prompt for your thinking rather than an answer.

Is its output audit evidence?

No. It is advisory analytical output for qualified professionals. Your firm's quality-management framework governs the engagement, and the audit is performed by your team.

05 · Pricing & efficiency

What the efficiency is actually worth, and how it is priced.

How is the service priced?

Under an Order Form, per engagement — and the Order Form is the authority on how usage is measured. The unit of measure and what a run includes are stated in your Order Form; ask us for the current language rather than relying on this page, which is not a price list. What we will say here is the design intent: the pricing is being shaped so that submitting the whole binder is not penalised against submitting one file at a time, because one file at a time is the usage that produces the worst results.

Will this reduce what our audits cost?

It changes where review hours go rather than eliminating them. The work it compresses is first-pass review: reading the binder end to end, checking each procedure was evidenced, spotting the tie-out that does not tie. What it does not compress is the judgment — assessing whether a response is adequate, deciding materiality, forming the conclusion. Whether that shows up as fewer hours, earlier issue detection, or more review depth for the same hours is a decision for your engagement planning, not something the tool decides for you.

Should we bill it to clients, or absorb it?

That is your firm's decision, and we do not have a view we would push. What we can give you is the factual basis for either conversation — what was reviewed, what was found, and what it cost.

06 · Talking to your own audit clients

Your audit clients, their audit committees, and peer reviewers will ask you the same questions you asked us. This section is what we can put in your hands for those conversations. It is factual input to your firm's own communications — not professional-standards advice, and not a substitute for your own quality-management framework or counsel.

Do we have to tell our audit clients we used it?

That is a determination for your firm, under your own standards, engagement letters, and client terms. What we can tell you is what there is to disclose: an AI-assisted analytical review of completed workpapers, performed inside your firm on your firm's own files, producing advisory review comments that your team dispositioned by name. No client data leaves your engagement to become anyone's training data, and the tool signs nothing. The decision is yours; we have no basis for telling you how other firms' disclosure conversations have gone, and we would rather say so than imply a pattern.

Our client asks: "did AI make decisions about our audit?"

"No. We used an AI-assisted review tool on our own workpapers, the way we would use any review checklist or analytical software. It raises questions for us; our people answer them, and every answer is recorded against a named person. No conclusion in your audit was made by software."

That statement is supported by mechanisms you can point to if pressed: mandatory human disposition on every item, named sign-off, and an exportable decision trail.

Our client asks: "was our confidential information exposed to an AI company?"

"Workpaper excerpts are processed by enterprise model providers through inference APIs, under terms that commit them not to train on our data. Our vendor holds those commitments, publishes its data-handling policy, and will produce the confirmations for our file. The vendor deletes the uploaded files when each run finishes; what it keeps is the result set and an archived copy of the workpapers we submitted, scoped to our firm, for the life of the engagement plus sixty days — so a re-check does not need a fresh upload — and it purges that on our request."

If your client asks for the underlying confirmations, ask us — we will provide them for your file rather than have you vouch for a vendor arrangement on our word alone.

If your client's own agreements restrict disclosure of their information to subprocessors, that is the constraint to check before you run their engagement through any tool, including ours. We will give you our subprocessor register for that check.

Our client wants to prohibit AI use on their engagement. Then what?

Tell us, and that engagement does not go through the service. We will record the restriction at the engagement level. Because our unit of work is an engagement binder rather than a firm-wide ingestion, honouring a per-client prohibition does not require you to change anything about your other engagements.

A peer reviewer or inspector asks how the tool was used and controlled.

We can supply the documentation set; you supply the engagement facts. On request we provide: this FAQ, the AI Use Policy, Capabilities & Known Limitations, and our security, privacy, vendor, incident-response and change-management policies. From the engagement itself you hold the coverage record (what was read), the finding set with dispositions and named sign-offs, and the exported review record. What we cannot supply — and no vendor can — is your firm's own documentation of how you evaluated the tool and integrated it into your quality-management framework.

Can we say our audits are "AI-powered" in marketing?

Please be careful, and please do not imply an outcome. No tool can guarantee a clean inspection or a particular regulatory result, and we do not make that claim for ourselves. Describing the tool as an AI-assisted review of workpapers performed by your team is accurate; describing your audits as AI-verified is not, and would misstate both what the tool does and where responsibility sits.

If a question is not answered here. Send it to us. If the honest answer is "not yet" or "we do not know", that is what we will say, and it will go on the Roadmap in the AI Use Policy with a date and an owner. We would rather lose a deal on a narrow answer than win one on a broad claim we cannot support in a vendor-risk review.