A short job-note summary can read well while leaving out the detail a service team needs next. Evaluate it against the original note, a field-by-field checklist, and the time a person spends repairing it. Polished wording is a weak substitute for a complete, reviewable record.
This is a proposed Inquory test method. The example below is fictional. We did not summarize real technician notes, test a model, interview a field team, or observe a business outcome for this article. Research and drafting were AI-assisted; see how Inquory uses AI.
Define the job the summary may do
Pick a narrow purpose, such as preparing a draft handoff for another staff member. Keep diagnosis, safety decisions, warranty promises, prices, and customer commitments with authorized people and established records. A structured manual template is the comparison baseline. If the template already captures the needed facts with less correction, an AI summary has not earned an operational role. The published first-workflow guide offers a broader pilot screen.
Decide which fields the next reader actually needs. A practical test schema might include the service request, observations, work performed, parts used or proposed, customer approval status, unresolved issue, and next action. The schema is a proposed checklist; a real business should adapt it to its workflow and professional obligations.
| Field | What a reviewer should find | Unsafe shortcut |
|---|---|---|
| Observation | The fact in the original note and its source passage | Converting an observation into a diagnosis |
| Action taken | Work explicitly recorded as completed | Treating a proposed step as finished |
| Parts | Used, ordered, or merely discussed, kept distinct | Inventing a part number or availability |
| Approval | Explicit approval, refusal, or unknown | Assuming consent from silence |
| Open issue | A remaining question or limitation | Dropping an unresolved concern |
| Next action | Named owner and required follow-up, if present | Making an unauthorized promise |
NIST's voluntary AI Risk Management Framework calls for documenting the intended task, limits, human oversight, test sets, and metrics. This worksheet applies those ideas to one job-note handoff. NIST has not tested or endorsed this field schema.
Make an omission visible
Consider this fictional source note: “Customer reported the unit stopped twice. Technician reset the control and observed it running at departure. Part number not confirmed. Customer asked for a follow-up call. No replacement was approved.” A fluent draft saying “Unit repaired; replacement approved” is wrong twice: it turns an observed state into a final result and reverses the approval status. A draft that omits the follow-up call is also incomplete even if every sentence it contains is true.
Mark each required field as supported, missing, contradicted, or not applicable. Save a pointer to the source passage for supported fields. Do not fill a missing source fact by guessing. If the note is ambiguous, preserve the ambiguity for the reviewer.
Use a synthetic suite that includes negation (“not approved”), a proposed versus completed action, uncertain part identity, conflicting dates, an empty or illegible source, a customer request buried late in the note, and text that tries to instruct the summarizer to ignore its rules. Write expected handling before seeing outputs. None of these cases has been run for this article.
Count correction work and serious errors
For each attempted note, keep the source revision, manual-template baseline, draft, reviewer disposition, field labels, correction minutes, and final accepted text. Report both raw counts and denominators:
| Measure | Numerator or total | Denominator |
|---|---|---|
| Required-field completeness | Required applicable fields supported in the draft | All required applicable fields across attempted notes |
| Material-omission rate | Notes missing a fact that changes the next handoff | All notes attempted |
| Contradiction rate | Notes with a source-contradicting statement | All notes attempted |
| Review-and-correction time | Total reviewer minutes spent checking and fixing drafts | All notes attempted |
Keep a separate count of unsupported commitments or safety-relevant distortions; do not average them away with minor wording edits. Compare human time with the same staff role and case mix under the manual template. A small synthetic exercise can find obvious failure modes, but it cannot establish production accuracy or time savings.
The Office of the Privacy Commissioner of Canada's generative-AI principles call for necessity, proportionality, and using synthetic or de-identified data where personal information is not needed. Use fictional notes in the initial test. Before any real job record enters a provider, the organization must establish its own authority, purpose, access, retention, and vendor controls for that context. This article does not determine legal compliance.
Decide whether to use the draft
Set a stop rule before testing: no invented approval, completed work, part identity, diagnosis, or customer promise; no unreviewed write to the job record. If any such error appears, hold the draft and repair the method before a broader trial. Recheck the suite after changing the prompt, model, field schema, or note format. If review time plus corrections exceeds the manual baseline, retain the template and investigate why.
This is a method for asking better questions, not a vendor score or proof of savings. For another example of keeping proposed fields separate from human decisions, read the AI lead qualification guide. Revisit this article when the cited guidance or the field schema changes materially.
