Count review, corrections and recurring upkeep alongside drafting; keep setup separate and leave missing evidence visible.
Measure the finished work
A reply that takes seconds to generate can still take several minutes to check. If staff must find the source, correct a condition, and ask a colleague whether the answer is usable, the generation time tells you little about the effort saved.
Once you have selected a bounded AI pilot, measure the work through to a checked result. This guide provides an original observation worksheet for a small business testing staff-reviewed reply drafting. All numerical examples below are hypothetical. They illustrate the accounting method, not AI Empower client results or expected savings.
Work through a hypothetical week
-
Current method
Suppose a fictional team compares two matched sets of 20 enquiries. Every case reaches the same checked-draft standard. The current method takes 180 active person-minutes in total.
-
AI assisted handling
The AI-assisted set takes 30 minutes for preparation and drafting, 45 for verification, 35 for corrections and manual reconstruction, and 20 for handoff preparation. Total effort is 130 minutes. Four rejected drafts are finished manually; their abandoned attempts and replacement work are already included in those categories.
-
Observed handling difference
The observed difference in this invented example is 50 minutes across 20 cases. Comparing only the 30 drafting minutes with the full 180-minute baseline would misstate the result.
-
Recurring work and setup
Now suppose the team also spends 40 minutes that week on additional source upkeep and routine checks required by the pilot. Subtracting that recurring effort leaves a ten-person-minute reduction in recurring effort, before setup. A separate three-hour setup exercise remains an upfront cost. If setup happened that week, total first-week effort increased by 170 person-minutes. Usable capacity has not yet been established.
These figures demonstrate arithmetic only. Twenty invented cases provide no evidence about another business, future demand, or dependable savings. A real comparison needs observations across the routine and difficult work the team actually receives.
Week 1 · Fictional local practice
Count the whole task
Compare one observation period and the same checked result. Include all staff, rejected attempts and manual completion. Totals include effort already spent on unresolved work.
Handling, recurring and first-period differences in person-minutes, with comparison warnings.
Use fictional, non-sensitive values only. Entries remain in page memory. No uploads, storage or network requests are made by this activity. Copy and print occur only when selected.
Use one shared observation period. One person working for five minutes is five person-minutes; two people working together for five minutes is ten. Do not enter elapsed waiting time.
Use the text-only exercise
The complete unchanged manuscript remains readable if the controls are unavailable.
Return to Work through a hypothetical weekWork the comparison on paper
Use the fictional figures above to make three separate calculations: handling difference, recurring-effort difference and first-period difference. Record the finish line and both completed-case counts beside the result. For this example, handling effort per completed case is 9 minutes for the current method and 6.5 minutes for the AI-assisted method. Including the additional weekly upkeep changes the assisted average to 8.5 minutes. These are calculated examples, not observations.
For your own sheet, leave a field marked unknown when it has not been measured. A blank is not zero. Keep effort already spent on unresolved cases in the handling total and also show that subset separately, without adding it twice. Do not draw a completed-case comparison when the methods used different quality standards or materially different case mixes.
Write one decision note containing the period, completed counts, unresolved work, handling totals, additional upkeep and separate setup hours. Keep negative savings visible. If the interactive effort worksheet is available on this page, use “Calculate effort differences” with fictional values. The paper instructions and row fields below work without it; the arithmetic does not establish cash savings or usable capacity.
Set a fair comparison
Choose one enquiry type and one finish line, such as “a checked reply ready for the authorised employee to send.” Use that finish line for both the current method and the AI-assisted method. Neither trial needs to send a customer message.
Write down the required facts, restrictions, and unanswered questions before timing either method. The same standard must apply to both. A draft missing an important restriction has not finished merely because it reads well.
The Canadian Centre for Cyber Security cautions that generated output can be incorrect or overlook relevant factors and recommends checking it against credible sources. Include that verification in the task rather than treating it as optional overhead.
Use comparable examples with similar complexity and the same approved sources. Rotate the order of methods or use fresh matched examples to reduce the advantage of already knowing the answer. Record differences in staff experience or interruptions that might explain a result.
Copy this effort worksheet
Create one row per enquiry per method, including all retries and manual fallback, with these fields:
-
Case code and enquiry category, without customer identifiers
-
Method: current process or AI-assisted process
-
Source and instruction versions; date of observation
-
Input preparation and drafting: active person-minutes
-
Source checking and review: active person-minutes
-
Corrections, retries, or manual reconstruction: active person-minutes
-
Handoff preparation: active person-minutes
-
Total active effort: the sum of the four effort fields
-
Elapsed duration from agreed start to finish, recorded separately
-
Result: accepted, completed manually, or unresolved
-
Quality findings: missing facts, unsupported additions, or other defects
-
Context and next action: what affected the case and who will investigate
Choose a single category for each minute of work. If a reviewer rewrites a sentence while checking it, split the observation reasonably or use one category consistently; do not count the same minute twice. Add effort from every employee involved. Two people working together for five minutes consume ten person-minutes, even though only five minutes pass.
Pause active-effort timing during unrelated work or waiting. Keep elapsed duration separate, because a smaller labour total does not establish a faster customer response.
Keep failed attempts in the record
If staff reject an AI draft and complete the reply manually, include the abandoned attempt and the subsequent manual work. Mark the result “completed manually.” Counting only accepted AI drafts would hide effort the workflow actually required.
Leave unresolved cases visible. Report their number and accumulated effort alongside completed cases; do not quietly remove them from the comparison. If the two methods finish different numbers or kinds of cases, the totals alone cannot establish which is more efficient.
For initial practice, invent enquiries and business facts. Canadian privacy regulators’ principles recommend synthetic or de-identified information when personal information is unnecessary. Real observations should use an approved internal process and capture only the information needed. A timing exercise does not require copying customer messages into a new AI tool.
Report the result without overstating it
Keep three lines in the decision note: observed handling-effort difference; additional recurring effort; and one-time setup effort. Record tool charges separately in money. Do not add dollars and minutes or treat unused salaried time as an automatic reduction in cash spending.
Name the useful work any released capacity could support, then check whether it actually becomes available. Ten scattered minutes may be harder to use than a predictable block of time. Quality still matters: a lower average does not justify unsupported promises or missed restrictions.
If correction time dominates, inspect the specific errors before changing the prompt. If source lookup dominates, improve access to the approved material. Repeat the comparison after one documented change, keeping previous failures visible.
Start with a small observation sheet and an honest result, including “no useful improvement yet.”
Sources rechecked 4 October 2026
A practical next step
Bring one task and a sample effort log.
Related reading: Choose a bounded first AI project.
Sources & review
Primary sources checked on . The checklists and planning examples are AI-assisted editorial guidance, not source quotations or reported client results.
- Canadian Centre for Cyber Security: Generative artificial intelligence
- Canadian privacy regulators: Generative AI principles
This AI-assisted guide uses fictional examples for practice. It does not report client results or establish that a live system will behave the same way.
Originally published: .
