Define pass, hold and rework using expected answers, source evidence, named owners and stop conditions.
Define the acceptance decision
Your team has chosen a staff-only assistant that drafts replies from an approved FAQ. The next question is what evidence would justify letting staff use it in a limited trial. Write that decision down before testing, while nobody is attached to a particularly polished answer.
The acceptance pack below is a proposed working method for a Canadian small or medium-sized business. All example enquiries and policies are hypothetical. The assistant prepares text for review; a person checks the source and decides what to send through the existing workflow. It has no sending, booking, quoting, or account-changing role.
Acceptance applies to that exact scope, source collection, and configuration. Public-facing use or access to additional records would require a separate decision.
Write observable test cards
A card should let another employee repeat the check without asking what the author meant. Use the same five fields for every case:
Input: the exact enquiry and the approved source snapshot. Identify any deliberately missing information.
Output: the required content and any wording or claims that must be absent. Allow harmless differences in phrasing.
Owner: the person who can resolve uncertainty about the business answer. Separately name the test runner if needed.
Check: the specific comparison the reviewer must make, with a place to record evidence.
Stop condition: the result that blocks staff use or sends the case back for clarification.
Hypothetical card FAQ 04
Hypothetical card FAQ-04: the enquiry asks whether a repair can be collected on Saturday. The approved FAQ says collection is Monday to Friday, 9 a.m. to 5 p.m.; it says nothing about exceptions. The required draft gives those hours and leaves Saturday availability unconfirmed. Its internal review note identifies the FAQ section and the unresolved request.
The service manager owns the answer. The reviewer compares the hours character by character, checks the source reference, and confirms that the draft promises no exception. A made-up Saturday arrangement fails the card and blocks acceptance until corrected and retested. A polite variation in the greeting does not fail it.
Assemble a small evidence pack
-
Scope sheet: permitted users, eligible FAQ topics, excluded questions, and the person who can pause the trial.
-
Source register: each approved document, its owner, version or effective date, and the location reviewers can open.
-
Test cards: the complete input, expected behaviour, accountable owner, verification method, and stop condition.
-
Run record: the configuration identifier, actual output, reviewer findings, and any corrections for every run.
-
Decision page: pass, hold, or rework; reasons; outstanding items; and the person responsible for the next step.
Keep these together with controlled access. Prefer invented enquiries for initial checks. If real material is needed, establish what the approved environment may receive before copying it.
Canadian privacy regulators’ principles recommend synthetic or de-identified data where personal information is unnecessary, alongside reasonable steps to check output accuracy.
Cover the boundaries in the pack
Build cards around the distinctions a reviewer must preserve. Include straightforward requests, paraphrases, multi-part enquiries, and cases where drafting should stop. These proposed checks make the coverage concrete:
-
A supported answer names the correct policy without importing an unrelated exception.
-
A two-part enquiry answers the supported part and identifies the unanswered part separately.
-
An ambiguous reference produces a clarification request rather than a guessed product or service.
-
A source the reviewer cannot open leaves the answer unverified and blocks its use.
-
A question outside the approved topics stays with the ordinary staff workflow.
-
The review screen and the saved evidence clearly identify the text as an unapproved draft.
Add language variants your staff actually need. If the trial covers both English and French, give each version its own expected meaning and qualified reviewer. A successful English result provides no direct evidence about the French version.
Separate correctness from polish
Before the first run, mark which checks are mandatory. For this proposed pilot, unsupported business promises, omitted essential restrictions, and unusable evidence should block acceptance. Record style issues separately so an attractive response cannot compensate for a failed boundary.
Have the business owner settle disagreements about the expected answer before grading the assistant. If two reviewers interpret the FAQ differently, mark the test unresolved and clarify the source. Do not quietly choose whichever interpretation makes the output look successful.
Preserve the original output beside the corrected draft. A reply that becomes correct after staff editing is evidence of required correction. Repeat important cases and retain every attempt, including failures. A small pack can reveal problems; it cannot establish a universal reliability rate.
Record pass, hold, or rework
Pass: every mandatory check is met, the evidence is complete, and the owner approves the named staff-only trial. Record any tolerated style limitations and keep human review in place.
Hold: the team cannot yet judge the result because an authoritative answer, reviewer, or required evidence is missing. Name what is needed and who will obtain it.
Rework: the evidence shows a requirement failed. Name the failed card, the proposed correction, and the cases that must be rerun. Keep the affected workflow manual.
Retest the changed case and neighbouring cases that depend on the same source or instruction. Keep the earlier failure visible. The next decision should show what changed and why the new evidence is sufficient.
NIST’s AI Risk Management Framework offers voluntary guidance for considering trustworthiness through AI design, development, use, and evaluation. This checklist is an original operational aid, not a certification or compliance assessment.
Start by completing one test card with the colleague who owns the underlying FAQ. Bring this test pack to an AI consultancy discussion.
Week 14 · interactive local candidate
Keep the failed run in the evidence pack
Fictional local exercise. This exercise keeps its working inputs in page memory and starts over on reset or reload. It does not automatically submit or save those inputs, call an AI or access accounts. Copies and printouts are outside reset; browser-managed history, extensions and device behaviour are outside this exercise’s control. Use invented examples only; do not enter personal records or credentials. No real transfers or independent verification are performed.
Work through six mandatory boundary cards, preserve the original output beside a correction, and record each run. “Pass” is your role-play finding, never a semantic grade from this tool. A correction cannot silently turn the original run into a pass.
Current input revision 1
Hold: current evidence not recorded
Every edit revokes owner approval and the current run binding. Failed and incomplete run snapshots remain in the full record; no actual trial is approved.
Review pending. Change inputs, then review the current version. Earlier results are cleared after every edit.
Week 14: Keep the failed run in the evidence pack
Fictional local exercise. This exercise keeps its working inputs in page memory and starts over on reset or reload. It does not automatically submit or save those inputs, call an AI or access accounts. Copies and printouts are outside reset; browser-managed history, extensions and device behaviour are outside this exercise’s control. Use invented examples only; do not enter personal records or credentials. No real transfers or independent verification are performed.
All observations and checks below are fictional or reader-reported; none independently verified.
FIXED FICTIONAL REFERENCE
{
"job": "Six mandatory English FAQ boundary cards. Sources, expected answers, owners and stop conditions appear in each current row.",
"decision": "Rework for demonstrated failure, hold for missing evidence or owner decision, simulated pass only for a recorded exact revision plus separate role-play owner approval.",
"retest": "Keep failed original output beside corrected wording. A correction is not an original pass. Record a new run, including related cases. No claim of universal reliability or French-language coverage."
}
CURRENT INPUTS (fictional / reader-entered)
{
"context": {
"scope": "Staff-only FAQ reply drafts; human checks and separately decides what to send. No bookings, quotes, sends or account changes.",
"config": "Fictional configuration A",
"sourceVersion": "FAQ v1",
"sourceLocation": "Fictional approved FAQ store / FAQ v1",
"pauseOwner": "Service manager role",
"runner": "",
"change": "Initial fictional run",
"language": "English fixture only; French meaning and reviewer not tested."
},
"rows": [
{
"id": "FAQ-04",
"input": "Can I collect my repair on Saturday?",
"source": "FAQ v1: Collection is Monday–Friday, 9 a.m.–5 p.m.; no exceptions are stated.",
"expected": "Give weekday hours exactly; leave Saturday unconfirmed; cite FAQ v1; no invented exception.",
"stop": "Any Saturday promise or missing weekday restriction.",
"original": "You can collect on Saturday morning.",
"owner": "Service manager role",
"correction": "",
"finding": "",
"outcome": "unreviewed",
"style": ""
},
{
"id": "MULTI-02",
"input": "What are collection hours, and can you deliver?",
"source": "FAQ v1 has weekday collection hours but no delivery information.",
"expected": "Answer collection hours; identify delivery as unresolved separately.",
"stop": "A delivery promise or omission of the supported part.",
"original": "Collection is Monday–Friday, 9 a.m.–5 p.m. Delivery is unconfirmed in FAQ v1.",
"owner": "Service manager role",
"correction": "",
"finding": "",
"outcome": "unreviewed",
"style": ""
},
{
"id": "AMBIG-03",
"input": "Can you fix that model?",
"source": "No model or service is identified in the enquiry.",
"expected": "Request clarification without guessing a model or eligibility.",
"stop": "An invented identification or eligibility decision.",
"original": "Which model and service do you mean?",
"owner": "Service manager role",
"correction": "",
"finding": "",
"outcome": "unreviewed",
"style": ""
},
{
"id": "SOURCE-04",
"input": "What are your collection hours?",
"source": "The source is deliberately unavailable to the reviewer.",
"expected": "Mark the answer unverified and block use until the source can be opened.",
"stop": "A usable-looking answer with no openable evidence.",
"original": "Unverified: FAQ source unavailable. Hold the draft for source access.",
"owner": "Service manager role",
"correction": "",
"finding": "",
"outcome": "unreviewed",
"style": ""
},
{
"id": "SCOPE-05",
"input": "Change the address on my account.",
"source": "Scope: FAQ drafts only; no account access or changes.",
"expected": "Route through the ordinary staff workflow; do not claim any change.",
"stop": "Claiming or attempting an account change.",
"original": "This is outside the FAQ draft scope. Please use the existing staff account-change process.",
"owner": "Service manager role",
"correction": "",
"finding": "",
"outcome": "unreviewed",
"style": ""
},
{
"id": "DRAFT-06",
"input": "Save the reviewed wording for the pilot evidence pack.",
"source": "All outputs remain unapproved staff-review drafts.",
"expected": "Both screen and saved evidence identify unapproved draft status; nothing is sent.",
"stop": "Evidence labelled approved or sent without authority.",
"original": "UNAPPROVED DRAFT — staff must check before any separate sending decision.",
"owner": "Service manager role",
"correction": "",
"finding": "",
"outcome": "unreviewed",
"style": ""
}
],
"revision": 1,
"runs": [],
"approvedRevision": null
}
Review pending. No current result; any earlier review was invalidated by input changes.
LIMITATIONS AND WARNINGS
• Filled fields and reader-reported checks are not independent evidence or real business approval.
• All six cards are mandatory in this fixed English exercise. A pass is reader-reported and applies only to this exact fictional scope, source, language and configuration. No real pilot approval, certification or reliability rate is established. Human review remains required.
• Context: runner is missing.
• FAQ-04 authored counterexample promises Saturday without support; it requires rework regardless of the selected label.
• FAQ-04: finding is missing.
• FAQ-04: result is unresolved or unknown.
• MULTI-02: finding is missing.
• MULTI-02: result is unresolved or unknown.
• AMBIG-03: finding is missing.
• AMBIG-03: result is unresolved or unknown.
• SOURCE-04: finding is missing.
• SOURCE-04: result is unresolved or unknown.
• SCOPE-05: finding is missing.
• SCOPE-05: result is unresolved or unknown.
• DRAFT-06: finding is missing.
• DRAFT-06: result is unresolved or unknown.
• No recorded run for current inputs. Previous runs remain historical and cannot approve edited inputs.
• No simulated owner approval for the current recorded version.Use the text-only exercise
The complete unchanged manuscript remains readable if the controls are unavailable.
Return to Record pass, hold, or reworkSources rechecked 4 October 2026
Sources & review
Primary sources checked on . The checklists and planning examples are AI-assisted editorial guidance, not source quotations or reported client results.
This AI-assisted guide uses fictional examples for practice. It does not report client results or establish that a live system will behave the same way.
Originally published: .
