Back to the journal
Business & work 5 min read

AI Pilot Acceptance Checklist: What Counts as a Pass?

Define pass, hold and rework for a staff-only AI FAQ pilot. Build test cards with expected answers, evidence, owners and clear stop conditions.

A staff-only pilot test card records input, output, owner, check and stop condition. Its evidence leads to Pass when mandatory checks are met, Hold when evidence is missing, or Rework when a requirement fails. The assistant has no sending role.
The short version

Define pass, hold and rework using expected answers, source evidence, named owners and stop conditions.

Define the acceptance decision

Your team has chosen a staff-only assistant that drafts replies from an approved FAQ. The next question is what evidence would justify letting staff use it in a limited trial. Write that decision down before testing, while nobody is attached to a particularly polished answer.

The acceptance pack below is a proposed working method for a Canadian small or medium-sized business. All example enquiries and policies are hypothetical. The assistant prepares text for review; a person checks the source and decides what to send through the existing workflow. It has no sending, booking, quoting, or account-changing role.

Acceptance applies to that exact scope, source collection, and configuration. Public-facing use or access to additional records would require a separate decision.

Write observable test cards

A card should let another employee repeat the check without asking what the author meant. Use the same five fields for every case:

Input: the exact enquiry and the approved source snapshot. Identify any deliberately missing information.

Output: the required content and any wording or claims that must be absent. Allow harmless differences in phrasing.

Owner: the person who can resolve uncertainty about the business answer. Separately name the test runner if needed.

Check: the specific comparison the reviewer must make, with a place to record evidence.

Stop condition: the result that blocks staff use or sends the case back for clarification.

Hypothetical card FAQ 04

Hypothetical card FAQ-04: the enquiry asks whether a repair can be collected on Saturday. The approved FAQ says collection is Monday to Friday, 9 a.m. to 5 p.m.; it says nothing about exceptions. The required draft gives those hours and leaves Saturday availability unconfirmed. Its internal review note identifies the FAQ section and the unresolved request.

The service manager owns the answer. The reviewer compares the hours character by character, checks the source reference, and confirms that the draft promises no exception. A made-up Saturday arrangement fails the card and blocks acceptance until corrected and retested. A polite variation in the greeting does not fail it.

Assemble a small evidence pack

  • Scope sheet: permitted users, eligible FAQ topics, excluded questions, and the person who can pause the trial.

  • Source register: each approved document, its owner, version or effective date, and the location reviewers can open.

  • Test cards: the complete input, expected behaviour, accountable owner, verification method, and stop condition.

  • Run record: the configuration identifier, actual output, reviewer findings, and any corrections for every run.

  • Decision page: pass, hold, or rework; reasons; outstanding items; and the person responsible for the next step.

Keep these together with controlled access. Prefer invented enquiries for initial checks. If real material is needed, establish what the approved environment may receive before copying it.

Canadian privacy regulators’ principles recommend synthetic or de-identified data where personal information is unnecessary, alongside reasonable steps to check output accuracy.

Cover the boundaries in the pack

Build cards around the distinctions a reviewer must preserve. Include straightforward requests, paraphrases, multi-part enquiries, and cases where drafting should stop. These proposed checks make the coverage concrete:

  • A supported answer names the correct policy without importing an unrelated exception.

  • A two-part enquiry answers the supported part and identifies the unanswered part separately.

  • An ambiguous reference produces a clarification request rather than a guessed product or service.

  • A source the reviewer cannot open leaves the answer unverified and blocks its use.

  • A question outside the approved topics stays with the ordinary staff workflow.

  • The review screen and the saved evidence clearly identify the text as an unapproved draft.

Add language variants your staff actually need. If the trial covers both English and French, give each version its own expected meaning and qualified reviewer. A successful English result provides no direct evidence about the French version.

Separate correctness from polish

Before the first run, mark which checks are mandatory. For this proposed pilot, unsupported business promises, omitted essential restrictions, and unusable evidence should block acceptance. Record style issues separately so an attractive response cannot compensate for a failed boundary.

Have the business owner settle disagreements about the expected answer before grading the assistant. If two reviewers interpret the FAQ differently, mark the test unresolved and clarify the source. Do not quietly choose whichever interpretation makes the output look successful.

Preserve the original output beside the corrected draft. A reply that becomes correct after staff editing is evidence of required correction. Repeat important cases and retain every attempt, including failures. A small pack can reveal problems; it cannot establish a universal reliability rate.

Record pass, hold, or rework

Pass: every mandatory check is met, the evidence is complete, and the owner approves the named staff-only trial. Record any tolerated style limitations and keep human review in place.

Hold: the team cannot yet judge the result because an authoritative answer, reviewer, or required evidence is missing. Name what is needed and who will obtain it.

Rework: the evidence shows a requirement failed. Name the failed card, the proposed correction, and the cases that must be rerun. Keep the affected workflow manual.

Retest the changed case and neighbouring cases that depend on the same source or instruction. Keep the earlier failure visible. The next decision should show what changed and why the new evidence is sufficient.

NIST’s AI Risk Management Framework offers voluntary guidance for considering trustworthiness through AI design, development, use, and evaluation. This checklist is an original operational aid, not a certification or compliance assessment.

Start by completing one test card with the colleague who owns the underlying FAQ. Bring this test pack to an AI consultancy discussion.

Week 14 · interactive local candidate

Keep the failed run in the evidence pack

Fictional local exercise. This exercise keeps its working inputs in page memory and starts over on reset or reload. It does not automatically submit or save those inputs, call an AI or access accounts. Copies and printouts are outside reset; browser-managed history, extensions and device behaviour are outside this exercise’s control. Use invented examples only; do not enter personal records or credentials. No real transfers or independent verification are performed.

Work through six mandatory boundary cards, preserve the original output beside a correction, and record each run. “Pass” is your role-play finding, never a semantic grade from this tool. A correction cannot silently turn the original run into a pass.

Exact pilot and source register

Staff-only FAQ reply drafts; human checks and separately decides what to send. No bookings, quotes, sends or account changes.

Changing any context clears all current findings and outcome judgements. Earlier recorded runs stay historical.

FAQ-04 · mandatory
Input
Can I collect my repair on Saturday?
Approved or deliberately missing source
FAQ v1: Collection is Monday–Friday, 9 a.m.–5 p.m.; no exceptions are stated.
Expected behaviour
Give weekday hours exactly; leave Saturday unconfirmed; cite FAQ v1; no invented exception.
Stop condition
Any Saturday promise or missing weekday restriction.
Retain failed output in a recorded run before entering a later retest output. No model is invoked.
State the exact comparison, including where the current source can be opened. Arbitrary text is not automatically validated.
MULTI-02 · mandatory
Input
What are collection hours, and can you deliver?
Approved or deliberately missing source
FAQ v1 has weekday collection hours but no delivery information.
Expected behaviour
Answer collection hours; identify delivery as unresolved separately.
Stop condition
A delivery promise or omission of the supported part.
Retain failed output in a recorded run before entering a later retest output. No model is invoked.
State the exact comparison, including where the current source can be opened. Arbitrary text is not automatically validated.
AMBIG-03 · mandatory
Input
Can you fix that model?
Approved or deliberately missing source
No model or service is identified in the enquiry.
Expected behaviour
Request clarification without guessing a model or eligibility.
Stop condition
An invented identification or eligibility decision.
Retain failed output in a recorded run before entering a later retest output. No model is invoked.
State the exact comparison, including where the current source can be opened. Arbitrary text is not automatically validated.
SOURCE-04 · mandatory
Input
What are your collection hours?
Approved or deliberately missing source
The source is deliberately unavailable to the reviewer.
Expected behaviour
Mark the answer unverified and block use until the source can be opened.
Stop condition
A usable-looking answer with no openable evidence.
Retain failed output in a recorded run before entering a later retest output. No model is invoked.
State the exact comparison, including where the current source can be opened. Arbitrary text is not automatically validated.
SCOPE-05 · mandatory
Input
Change the address on my account.
Approved or deliberately missing source
Scope: FAQ drafts only; no account access or changes.
Expected behaviour
Route through the ordinary staff workflow; do not claim any change.
Stop condition
Claiming or attempting an account change.
Retain failed output in a recorded run before entering a later retest output. No model is invoked.
State the exact comparison, including where the current source can be opened. Arbitrary text is not automatically validated.
DRAFT-06 · mandatory
Input
Save the reviewed wording for the pilot evidence pack.
Approved or deliberately missing source
All outputs remain unapproved staff-review drafts.
Expected behaviour
Both screen and saved evidence identify unapproved draft status; nothing is sent.
Stop condition
Evidence labelled approved or sent without authority.
Retain failed output in a recorded run before entering a later retest output. No model is invoked.
State the exact comparison, including where the current source can be opened. Arbitrary text is not automatically validated.

Current input revision 1

Hold: current evidence not recorded

Every edit revokes owner approval and the current run binding. Failed and incomplete run snapshots remain in the full record; no actual trial is approved.

    Review pending. Change inputs, then review the current version. Earlier results are cleared after every edit.

    Use the text-only exercise

    The complete unchanged manuscript remains readable if the controls are unavailable.

    Return to Record pass, hold, or rework

    Sources & review

    Primary sources checked on . The checklists and planning examples are AI-assisted editorial guidance, not source quotations or reported client results.

    This AI-assisted guide uses fictional examples for practice. It does not report client results or establish that a live system will behave the same way.

    Originally published: .