Before an AI workflow reaches customers or records, review its task, boundary, staff experience, failure handling, and fallback. A polished demo is not enough evidence that it is ready to operate.
This is a proposed Inquory review checklist. It does not certify a system, replace professional advice, or report hands-on testing. Research and drafting were AI-assisted; see how Inquory uses AI.
1. Review the task and output
Write the one job the workflow performs, the cases it may accept, and the decision it supports. Confirm that the output contains only the fields the next staff member needs. Label generated content as a proposal, preserve unknowns, and require a source passage for extracted facts.
Keep diagnosis, safety decisions, pricing, customer commitments, rejection, payment, and deletion outside the pilot unless each action has separate authority and testing.
2. Review the operating boundary
Name the exact input, proposed output, approving person, forbidden actions, manual fallback, and stop authority. Separate deterministic controls from generated text. Confirm that missing information stays missing and that an uncertain write is reconciled before retry.
| Review item | Evidence to retain |
|---|---|
| Intended task | One bounded purpose and eligible-case definition |
| Human decision | Named role and accept, correct, reject options |
| Forbidden action | Customer, record, payment, safety, and deletion limits |
| Failure handling | Stop condition, repair owner, and manual fallback |
| Measurement | Attempt denominator, serious errors, review time, and baseline |
| Change control | Model, prompt, field map, integration, and article revision |
NIST's voluntary AI Risk Management Framework Core recommends defined human-AI responsibilities, documented test sets and metrics, comparisons with benchmarks, and testing in conditions similar to deployment. Use those ideas to make the review repeatable; a checklist alone does not prove that a control works.
3. Review the staff experience
Ask a person who does the real task to use the review screen. They should be able to compare the proposal with its source, find conflicts, correct or reject fields, and see whether anything changed in another system. Show errors plainly and keep the original input available.
Use the keyboard and a narrow screen as part of the exercise. Check visible focus, zoom and reflow, field labels, status messages, table scrolling, and the path back to the queue. W3C's guidance for WCAG 2.2 headings and labels explains that headings and labels should describe their topic or purpose. Correct markup and understandable wording need separate checks.
Time the review and correction work. Record where staff hesitate, open another system, or copy information manually. A faster generated draft can still make the whole task slower if comparison and repair are difficult.
4. Run synthetic acceptance cases
Create fictional cases for missing information, negation, conflicting statements, repeated delivery, an out-of-scope request, a timeout after a write, and text that tells the system to ignore its instructions. Write the expected proposal, reviewer action, and forbidden action first.
Retain every attempt and its final disposition. Count unsupported facts, material omissions, incorrect routing, unknown outcomes, review minutes, correction minutes, and forbidden actions. A passing average must not conceal a customer message, record change, or other serious action outside the boundary.
5. Sign off a limited rollout
Use one of four outcomes: proceed to a limited rollout, revise and retest, stop, or unknown. “Unknown” is appropriate when the evidence cannot establish what happened. Require signoff from the workflow owner, the staff role doing the review, and the person responsible for access and recovery. Record unresolved concerns rather than converting them into future tasks after launch.
For a limited rollout, cap the eligible cases and duration. Keep human approval on every output, monitor errors and repair time, and name the person who can stop the workflow. Verify the manual fallback before starting. If the workflow touches a CRM, use the CRM handoff checklist for duplicate delivery, stale updates, permissions, and destination readback.
Review the first cases promptly and stop on a forbidden action, unresolved unknown outcome, loss of source traceability, or unavailable fallback. Re-run affected checks when the model, prompt, workflow, permissions, data, interface, or integration changes. The result is a bounded operating decision, not a guarantee that later cases will behave the same way.
