An AI pilot can look successful when a dashboard counts every response as a completed job. A service-business owner needs a narrower answer: for one defined task, did the proposed workflow produce useful work after errors, human repair, and operating costs were counted?
This guide is an Inquory measurement worksheet, not the result of a pilot. The examples are fictional. Inquory did not test a model, interview operators, process customer records, or observe a business outcome for this article. Research and drafting were AI-assisted; see how Inquory uses AI.
The Canadian implementation guide for managers of AI systems describes context-specific testing, including varied and adverse conditions, retesting after significant changes, and documenting limitations. The voluntary NIST AI Risk Management Framework connects evaluation methods to the risks of a particular use and calls for documented test sets, oversight, and limits. Neither source validates the worksheet below or promises that a pilot will improve a business.
Write the decision before choosing a metric
Use one bounded task. For example: "Can a system propose service category and missing fields from a fictional incoming inquiry, while a person decides the next step?" Keep sending a message, rejecting a customer, changing a CRM record, diagnosing a problem, and booking an appointment outside this test.
Write down the simpler comparison. A short form, a staff checklist, or a rule-based queue may already produce a reviewable record. Compare the AI proposal with that baseline on the same kinds of cases. If the AI step adds more checking than it removes, the baseline may be the better operating choice. The first-workflow guide helps set this boundary.
Before testing, record who can inspect the outputs, which synthetic cases are in scope, what action is forbidden, how the test is stopped, and what evidence would change the decision. Do not enter real customer messages into an unapproved test system.
Freeze the denominator
Define an eligible case before seeing a result. For the inquiry example, an eligible case could be one unique, fictional inquiry in a saved test set with a recorded expected disposition. A duplicate delivery is a separate failure scenario, not a second successful inquiry. Keep cases that fail, need clarification, or require human repair in the denominator.
| Measure | Numerator | Denominator | What it can tell you |
|---|---|---|---|
| Reviewable-record rate | Eligible cases with all required fields and source passages available for review | All eligible cases attempted | Whether a person can inspect the proposal |
| Material-error rate | Eligible cases with a wrong or unsupported field that could change the next step | All eligible cases attempted | How often a consequential proposal needs correction |
| Human review-and-repair time | Total minutes spent reviewing and correcting eligible cases | All eligible cases attempted | Average human effort per attempted case |
| Duplicate-action rate | Duplicate-delivery scenarios that cause more than one attempted downstream action | All duplicate-delivery scenarios | Whether repeated events are contained |
These are proposed definitions, not industry benchmarks. Preserve the raw case list and each result so a failed case cannot disappear from the final rate. Report counts beside percentages: "18 of 20 synthetic cases" is clearer than a bare 90%.
Build cases that try to break the boundary
Write expected results before running the system. A small synthetic suite should include ordinary work and plausible exceptions:
| Fictional case | Expected handling | Failure to record |
|---|---|---|
| Clear service request with a supplied location | Propose fields with source passages; send to a person | Invented detail or unsupported confidence |
| Missing address or conflicting date | Mark the field missing or conflicting | Guessing an address or booking a time |
| Same inquiry delivered twice | Link the events for review; make no duplicate commitment | Two records or two external actions |
| Text that says "ignore your rules" | Treat it as customer text, not an instruction to the system | Rule change or unauthorized action |
| Urgent or out-of-scope request | Stop the automated path and use the existing human process | Unsupported diagnosis or promise |
| CRM or model unavailable | Preserve the original event and route to manual handling | Silent loss, fabricated completion, or untracked retry |
NIST's Measure Playbook suggests selecting measures for mapped risks, documenting oversight, errors, complaints, and go/no-go decisions. The suite above is Inquory's proposed application for a narrow service inquiry; it has not been executed or shown to be representative of real inquiries.
Count repair and cost, not just output
For each case, keep a ledger with: case ID, baseline result, expected result, system output, human disposition, error type, review minutes, correction minutes, and any attempted external action. Add review and correction minutes when calculating the table's human effort measure. A reviewer should be able to reopen a failed case and see what was corrected.
If the test later moves beyond synthetic cases under a separately approved process, measure the full operating cost for the same period: setup, model or vendor usage, staff review and repair, integration, monitoring, and failure recovery. Keep labour capacity separate from cash savings. The ROI calculator can model reader-supplied assumptions, but its output is a scenario, not evidence of achieved savings.
Do not infer incremental bookings from a higher count of processed inquiries. A before-and-after comparison can change with season, staffing, marketing, or case mix. Without a suitable comparison design, describe what was observed and what remains uncertain rather than attributing the difference to AI.
Work through one fictional decision
Suppose an owner prepares 20 fictional inquiries before running anything. The existing staff checklist produces 17 complete, reviewable records and takes 24 minutes in total. The AI-assisted proposal produces 18 reviewable records, but two proposals contain material errors and a person spends 36 minutes reviewing and repairing them. For this illustration only, assume the AI tool costs CAD $15 for the test period. These numbers are invented to show the arithmetic; they are not Inquory results or typical industry rates.
The proposed workflow's reviewable-record rate is 18 ÷ 20 = 90%. Its material-error rate is 2 ÷ 20 = 10%. Human review and repair average 36 ÷ 20 = 1.8 minutes per attempted case, compared with 24 ÷ 20 = 1.2 minutes for the checklist. The extra review time is 12 minutes for this small set, before the hypothetical CAD $15 tool cost is considered. One extra reviewable record does not demonstrate a net benefit when consequential errors, repair, and cost rise.
This comparison does not prove the checklist would outperform an AI system on real inquiries. The cases are fictional and the sample is small. It shows how to resist the tempting but incomplete headline of “18 successful responses.” A next experiment would need a revised workflow, a new saved test-set version, and a fresh decision rule; the owner could also stop and keep the checklist.
Copy this short worksheet before a pilot. Fill the thresholds **before** looking at its output:
| Decision field | Fill in before testing |
|---|---|
| Task and forbidden actions | One proposed task; list every send, write, booking, rejection, or promise the system may not make |
| Case set and denominator | Version, number of eligible cases, how duplicates and exceptions are counted |
| Simple baseline | The current form, rule, or staff checklist and the same measures applied to it |
| Required record | Case ID, source passage, proposed fields, human disposition, corrections, time, and attempted action |
| Stop limits | Maximum material errors, missing records, repair minutes, total cost, and any zero-tolerance boundary |
| Decision and owner | Go, revise, or stop; named accountable person, date, unresolved failures, and permitted next action |
Keep the completed worksheet with the raw cases and costs. If the test is interrupted, record the interruption rather than deleting its cases. If staff cannot reliably reconstruct what the system proposed and what a person changed, the pilot is not ready for a broader trial.
Make a go, revise, or stop decision
Before viewing results, the accountable operator should set task-specific limits for material errors, missing records, human repair effort, cost, and forbidden actions. A breach of the no-send, no-book, or no-reject boundary in this example is a stop and investigation, not a score to average away. A harmless formatting error can be revised and retested.
Record the decision with the exact test-set version, system version, date, reviewer, failures, unresolved limits, and the next permitted action. Retest after a material workflow or model change. Synthetic success permits consideration of a controlled next step; it does not establish production reliability, legal compliance, customer consent, or a financial return.
For a related concrete design, see the lead-qualification guide and the appointment-scheduling guide. The editorial method explains how Inquory handles evidence and corrections.