Accounting AI Pilot Acceptance: A Go/Hold/Stop Record

Accept an accounting AI pilot only for a named workflow, version and operating scope, using a signed go/hold/stop record. The record should connect agreed criteria to observed evidence, unresolved exceptions, retained human decisions and an exercised stop procedure. A good demonstration or readiness score does not replace that decision.
This is an original Fixed Labs concept framework, not a tested customer acceptance procedure, industry benchmark or compliance certification. The blank record is usable as a starting point; the worked example is entirely synthetic. Adapt the criteria with the people responsible for the work, data and professional obligations.
Decide on the next bounded use—not “AI for the firm”

An acceptance decision answers a narrow question: may this version of this workflow proceed under these conditions?
For example, “prepare unsent invoice drafts for the approved queue, with a billing reviewer checking every draft” describes a permission worth evaluating. “Automate billing” does not say whether the system may change approved terms, send an invoice or handle a disputed account.
The CPA AI readiness scorecard addresses an earlier decision: which workflows have sufficient documentation, inputs, controls and measurable value to consider a pilot? This acceptance record addresses what the pilot evidence now justifies. It does not repeat the scorecard or turn its score into production permission.
There is a second distinction. A pilot acceptance record is the sponsor’s decision about a bounded deployment. The operational approval map determines who may authorize a particular action during a run. Passing the pilot does not grant the system authority over every later transaction. For a concrete execution example, the published Susan billing workflow describes invoice preparation and human review; its customer results are not evidence that your pilot has passed.
Define go, hold and stop before seeing the results

Use these as separate dispositions, not levels on an autonomy ladder.
| Decision | Meaning | What happens next |
|---|---|---|
| Go | The evidence supports the stated next scope, with its controls and accepted residual risks | Start only that scope, until its review date or stop trigger |
| Hold | Evidence is missing or a remediable criterion has failed | Do not expand permission; assign repairs and a retest |
| Stop | An unacceptable risk, boundary breach or unavailable control makes continuation inappropriate | Suspend the affected activity, preserve evidence and move to the agreed fallback |
A hold must specify what remains permitted. It may allow synthetic evaluation while forbidding live access. A stop must specify what stops: one queue, one integration or the entire pilot. Neither label should leave staff guessing.
NIST’s voluntary AI RMF Playbook provides useful foundations without prescribing this form. MANAGE 1.1 asks whether the system meets its intended purpose and whether development or deployment should proceed. MEASURE 1.1 addresses measurement approaches and acceptable performance limits. MANAGE 2.4 addresses responsibilities and mechanisms to supersede, disengage or deactivate a system. This article turns those ideas into a proposed accounting-workflow decision record, not a NIST-approved checklist.
Build an evidence packet that a reviewer can inspect

Agree the criteria before the evaluation. Record the procedure and configuration version, the workload represented, the expected result for each case and the reviewer’s method. If criteria change after results arrive, record the change and repeat the affected evaluation rather than silently relabeling a failure.
Include ordinary cases and the exceptions that can change the decision: missing instructions, conflicting approvals, held work, an existing draft, an interrupted run and an input outside the agreed scope. Choose cases for the workflow’s risks; there is no universal sample size in this framework.
For every case, keep the original output, the reviewer’s correction and the saved final state separately. A corrected output can be acceptable for supervised work without proving the system produced it correctly on its own.
NIST MEASURE 2.1 calls for documenting test sets, metrics and evaluation tools; MEASURE 2.3 addresses evaluation under conditions similar to deployment. Use those principles to avoid treating a polished demonstration as representative operating evidence.
Separate quality, escalation and effort
Count each attempted case once in the case inventory. Then record the outcome categories and corrections explicitly. An appropriate escalation is not automatically an error; an incomplete case is not a verified completion; and a retry is not an additional invoice prepared.
If you report a rate, state the denominator, error definition, observation window and configuration version. Keep challenge-case results distinct from ordinary workload results. A deliberately exception-heavy evaluation set cannot establish the exception frequency of future live work.
For business value, compare equivalent work. Include review, correction and exception handling in staff effort. Do not compare automated draft preparation with the entire manual billing cycle or describe modeled time as measured savings. A go decision can be a learning decision without a proven ROI claim.
Copy this pilot acceptance decision record

Complete the record and attach its evidence before granting the next scope. Blank fields are unresolved decisions, not implied approvals. Keep evidence references in an access-controlled location; this template does not require publishing client records.
1. Decision identity and scope
| Field | Entry to complete |
|---|---|
| Record ID and revision | [Unique decision ID; revision; superseded record, if any] |
| Workflow and intended purpose | [Trigger, input, task and definition of done] |
| Evaluated version | [Procedure, rules, model/tool configuration and integration versions] |
| Evaluation window and environment | [Dates; synthetic, sandbox or authorized live setting] |
| Next permitted scope | [Queue, work types, systems, period, limits and review mode] |
| Explicitly excluded actions | [For example: sending, posting, changing terms or professional judgment] |
| Evidence packet | [Case inventory, expected results, observed outputs, corrections and final-state checks] |
| Decision expiry | [Date/time and event-based expiry triggers] |
2. Criteria and findings
Use one row per criterion. Add rows rather than collapsing unlike risks into an overall score.
| Criterion | Agreed pass condition | Evidence reference and finding | Disposition |
|---|---|---|---|
| Output quality | [Required fields and accounting checks; permitted corrections] | [Expected versus observed result, including failures] | [Pass/fail/not evaluated] |
| Exceptions | [Required escalations and forbidden guesses] | [Challenge cases and routing outcome] | [Pass/fail/not evaluated] |
| Scope and access | [Allowed data/actions and blocked actions] | [Observed boundary checks; limitations] | [Pass/fail/not evaluated] |
| Saved-state verification | [How to confirm what was actually saved] | [Readback or independent verification evidence] | [Pass/fail/not evaluated] |
| Interruption and recovery | [How to identify existing work and avoid repeat actions] | [Interrupted-run and recovery observations] | [Pass/fail/not evaluated] |
| Reviewer availability | [Named owner, coverage and response commitment] | [Coverage plan and observed handoff] | [Pass/fail/not evaluated] |
| Stop and fallback | [Who stops activity; safe manual continuation] | [Stop exercise, affected-work inventory and reconciliation] | [Pass/fail/not evaluated] |
| Operational value | [Comparable workload and effort measure] | [Observed result or explicit measurement gap] | [Pass/fail/not evaluated] |
Do not average a failed critical control against a strong value result. Mark which criteria are non-negotiable before the evaluation. A criterion not evaluated is not a pass.
3. Exceptions, obligations and decision
| Field | Entry to complete |
|---|---|
| Unresolved exceptions | [Issue, affected cases, severity and why unresolved] |
| Residual risk accepted | [Specific limitation, reason, scope and authorized risk owner] |
| Remediation | [Action, owner, due date, retest evidence and closure approver] |
| Stop triggers and response | [Trigger, immediate restriction, responder, notification and fallback] |
| Monitoring | [What is checked, by whom, when and where findings are recorded] |
| Final decision | [Go / Hold / Stop; rationale tied to criterion findings] |
| Permission retained during hold | [Exact activity still allowed, or none] |
| Conditions for reconsideration | [Required repairs, new evidence and affected cases to rerun] |
4. Signatures
- Workflow owner: [Name, role, signature/recorded approval, date and time]. Confirms the operating scope and definition of done.
- Evidence reviewer: [Name, role, signature/recorded approval, date and time]. Confirms which evidence was inspected and records limitations.
- Access/control owner: [Name, role, signature/recorded approval, date and time]. Confirms the stated restrictions and stop/fallback arrangements.
- Accountable sponsor: [Name, role, signature/recorded approval, date and time]. Records the final disposition and any residual risk they are authorized to accept.
One person may hold more than one role if your governance permits it; record that overlap. The system being evaluated cannot sign its own acceptance. A typed name without the firm’s agreed approval method is not evidence of a real signature.
Write stop conditions that someone can act on

Define the immediate restriction and the responsible person for every trigger. For an unsent invoice-drafting pilot, proposed triggers include:
- Outside-scope data or action: block further affected activity and have the access owner investigate. Do not retry with broader permissions.
- Missing or conflicting authority: stop that case and route it to the billing decision owner. Do not infer permission from a prior invoice.
- Saved result differs from the approved instruction: suspend the affected queue until the reviewer reconciles source, draft and correction evidence.
- Uncertain prior action after interruption: do not create another record until the operator determines whether the earlier action committed.
- Required logs, verification or reviewer coverage unavailable: hold the affected work; use the approved manual process.
- A material change to scope, rules, configuration or integration behavior: expire acceptance for the changed portion and rerun the relevant checks.
Stopping an automation is not the same as reversing its effects. The fallback must inventory affected work, reconcile saved records and route corrections through the authorized process. Do not promise to “roll back” an invoice already sent or a payment already initiated as though disabling a system undoes it.
NIST MANAGE 4.1 includes post-deployment monitoring, incident response, recovery and change management. A signed go therefore needs continuing review, not just an acceptance meeting.
Synthetic example: hold the live pilot, not the evidence

Everything in this example is invented for illustration. It is not a customer case, observed run or actual signature. The evidence labels below are fictional placeholders.
The proposed workflow prepares unsent invoice drafts from approved billing instructions. Sending, changing client terms and deciding tax treatment are excluded. The next requested scope is one approved queue with a reviewer checking every draft.
| Record field | Synthetic entry |
|---|---|
| Record and version | DEMO-ACCEPT-01; procedure v1; sandbox configuration A |
| Evaluated workload | A synthetic ordinary case plus missing-authority, held-time, existing-draft and interrupted-run challenge cases |
| Passing findings | The fictional case notes show held time excluded, missing authority escalated and saved draft fields checked |
| Blocking finding | The interrupted-run case lacks a conclusive saved-state check; a repeat creation could not be ruled out |
| Decision | Hold live drafting; synthetic sandbox evaluation only |
| Repair | Recovery owner adds an existing-work check and records the result of a repeated interruption exercise |
| Retest condition | Evidence reviewer inspects the recovery result and any affected draft checks before sponsor reconsideration |
| Stop condition | An uncertain prior action blocks further creation; operator reconciles the existing-work inventory |
| Expiry | Any change to procedure, configuration or requested scope requires reconsideration |
Illustrative signature entries—not real approvals: workflow owner [fictional role acknowledgment]; evidence reviewer [fictional evidence acknowledgment]; control owner [fictional restriction acknowledgment]; sponsor [fictional HOLD acknowledgment]. In an actual record, each would be replaced by a named person’s recorded approval with a date and time.
The ordinary case did not erase the recovery gap. The hold also did not require abandoning the workflow: it identified a repair, an owner and the evidence needed for a new decision. A later go would require a new signed revision; it would not retroactively make the original hold a pass.
Use the record at the acceptance meeting

Ask the sponsor to read three entries aloud: the next permitted scope, the unresolved risks and the stop response. If the people responsible disagree on any of them, resolve that disagreement before recording go.
For the preparation work that precedes evaluation, see Fixed Labs’ billing implementation account. It explains how documented rules and exceptions inform supervised work. This acceptance template is a separate proposed artifact—not a claim that the featured deployment used this exact form.
If your team needs help deciding what to evaluate first, start with the Fixed Labs AI Assessment. Use the acceptance record to organize the questions you want to resolve; an assessment is not certification or automatic permission to deploy.
Authoring note: Prepared with AI assistance under the owner’s unattended-article authorization, with Manuel Castillo named as author. This article does not claim a separate human editing or sign-off event. The decision framework is proposed; the worked example is synthetic.