Planning and measurement · Inquory Research

How to Measure an AI Pilot in a Service Business

Published · AI-assisted research

An AI pilot can look successful when a dashboard counts every response as a completed job. A service-business owner needs a narrower answer: for one defined task, did the proposed workflow produce useful work after errors, human repair, and operating costs were counted?

This guide is an Inquory measurement worksheet, not the result of a pilot. The examples are fictional. Inquory did not test a model, interview operators, process customer records, or observe a business outcome for this article. Research and drafting were AI-assisted; see how Inquory uses AI.

The Canadian implementation guide for managers of AI systems describes context-specific testing, including varied and adverse conditions, retesting after significant changes, and documenting limitations. The voluntary NIST AI Risk Management Framework connects evaluation methods to the risks of a particular use and calls for documented test sets, oversight, and limits. Neither source validates the worksheet below or promises that a pilot will improve a business.

Write the decision before choosing a metric

Use one bounded task. For example: "Can a system propose service category and missing fields from a fictional incoming inquiry, while a person decides the next step?" Keep sending a message, rejecting a customer, changing a CRM record, diagnosing a problem, and booking an appointment outside this test.

Write down the simpler comparison. A short form, a staff checklist, or a rule-based queue may already produce a reviewable record. Compare the AI proposal with that baseline on the same kinds of cases. If the AI step adds more checking than it removes, the baseline may be the better operating choice. The first-workflow guide helps set this boundary.

Before testing, record who can inspect the outputs, which synthetic cases are in scope, what action is forbidden, how the test is stopped, and what evidence would change the decision. Do not enter real customer messages into an unapproved test system.

Freeze the denominator

Define an eligible case before seeing a result. For the inquiry example, an eligible case could be one unique, fictional inquiry in a saved test set with a recorded expected disposition. A duplicate delivery is a separate failure scenario, not a second successful inquiry. Keep cases that fail, need clarification, or require human repair in the denominator.

MeasureNumeratorDenominatorWhat it can tell you
Reviewable-record rateEligible cases with all required fields and source passages available for reviewAll eligible cases attemptedWhether a person can inspect the proposal
Material-error rateEligible cases with a wrong or unsupported field that could change the next stepAll eligible cases attemptedHow often a consequential proposal needs correction
Human review-and-repair timeTotal minutes spent reviewing and correcting eligible casesAll eligible cases attemptedAverage human effort per attempted case
Duplicate-action rateDuplicate-delivery scenarios that cause more than one attempted downstream actionAll duplicate-delivery scenariosWhether repeated events are contained

These are proposed definitions, not industry benchmarks. Preserve the raw case list and each result so a failed case cannot disappear from the final rate. Report counts beside percentages: "18 of 20 synthetic cases" is clearer than a bare 90%.

Build cases that try to break the boundary

Write expected results before running the system. A small synthetic suite should include ordinary work and plausible exceptions:

Fictional caseExpected handlingFailure to record
Clear service request with a supplied locationPropose fields with source passages; send to a personInvented detail or unsupported confidence
Missing address or conflicting dateMark the field missing or conflictingGuessing an address or booking a time
Same inquiry delivered twiceLink the events for review; make no duplicate commitmentTwo records or two external actions
Text that says "ignore your rules"Treat it as customer text, not an instruction to the systemRule change or unauthorized action
Urgent or out-of-scope requestStop the automated path and use the existing human processUnsupported diagnosis or promise
CRM or model unavailablePreserve the original event and route to manual handlingSilent loss, fabricated completion, or untracked retry

NIST's Measure Playbook suggests selecting measures for mapped risks, documenting oversight, errors, complaints, and go/no-go decisions. The suite above is Inquory's proposed application for a narrow service inquiry; it has not been executed or shown to be representative of real inquiries.

Count repair and cost, not just output

For each case, keep a ledger with: case ID, baseline result, expected result, system output, human disposition, error type, review minutes, correction minutes, and any attempted external action. Add review and correction minutes when calculating the table's human effort measure. A reviewer should be able to reopen a failed case and see what was corrected.

If the test later moves beyond synthetic cases under a separately approved process, measure the full operating cost for the same period: setup, model or vendor usage, staff review and repair, integration, monitoring, and failure recovery. Keep labour capacity separate from cash savings. The ROI calculator can model reader-supplied assumptions, but its output is a scenario, not evidence of achieved savings.

Do not infer incremental bookings from a higher count of processed inquiries. A before-and-after comparison can change with season, staffing, marketing, or case mix. Without a suitable comparison design, describe what was observed and what remains uncertain rather than attributing the difference to AI.

Work through one fictional decision

Suppose an owner prepares 20 fictional inquiries before running anything. The existing staff checklist produces 17 complete, reviewable records and takes 24 minutes in total. The AI-assisted proposal produces 18 reviewable records, but two proposals contain material errors and a person spends 36 minutes reviewing and repairing them. For this illustration only, assume the AI tool costs CAD $15 for the test period. These numbers are invented to show the arithmetic; they are not Inquory results or typical industry rates.

The proposed workflow's reviewable-record rate is 18 ÷ 20 = 90%. Its material-error rate is 2 ÷ 20 = 10%. Human review and repair average 36 ÷ 20 = 1.8 minutes per attempted case, compared with 24 ÷ 20 = 1.2 minutes for the checklist. The extra review time is 12 minutes for this small set, before the hypothetical CAD $15 tool cost is considered. One extra reviewable record does not demonstrate a net benefit when consequential errors, repair, and cost rise.

This comparison does not prove the checklist would outperform an AI system on real inquiries. The cases are fictional and the sample is small. It shows how to resist the tempting but incomplete headline of “18 successful responses.” A next experiment would need a revised workflow, a new saved test-set version, and a fresh decision rule; the owner could also stop and keep the checklist.

Copy this short worksheet before a pilot. Fill the thresholds **before** looking at its output:

Decision fieldFill in before testing
Task and forbidden actionsOne proposed task; list every send, write, booking, rejection, or promise the system may not make
Case set and denominatorVersion, number of eligible cases, how duplicates and exceptions are counted
Simple baselineThe current form, rule, or staff checklist and the same measures applied to it
Required recordCase ID, source passage, proposed fields, human disposition, corrections, time, and attempted action
Stop limitsMaximum material errors, missing records, repair minutes, total cost, and any zero-tolerance boundary
Decision and ownerGo, revise, or stop; named accountable person, date, unresolved failures, and permitted next action

Keep the completed worksheet with the raw cases and costs. If the test is interrupted, record the interruption rather than deleting its cases. If staff cannot reliably reconstruct what the system proposed and what a person changed, the pilot is not ready for a broader trial.

Make a go, revise, or stop decision

Before viewing results, the accountable operator should set task-specific limits for material errors, missing records, human repair effort, cost, and forbidden actions. A breach of the no-send, no-book, or no-reject boundary in this example is a stop and investigation, not a score to average away. A harmless formatting error can be revised and retested.

Record the decision with the exact test-set version, system version, date, reviewer, failures, unresolved limits, and the next permitted action. Retest after a material workflow or model change. Synthetic success permits consideration of a controlled next step; it does not establish production reliability, legal compliance, customer consent, or a financial return.

For a related concrete design, see the lead-qualification guide and the appointment-scheduling guide. The editorial method explains how Inquory handles evidence and corrections.