Planning and measurement · Inquory Research

What to Record When an AI Workflow Changes Before Retesting

Published · AI-assisted research

A baseline binder and a workflow with one changed component lead to selected test cards, an unresolved card held for review, and a manual fallback notebook.
AI-generated editorial illustration · Inquory

Inquory's proposed change record connects the last accepted version, one recorded change, selected retests and a decision or manual fallback. The held card represents an unresolved outcome awaiting reconciliation, not an automatic retry. The shapes and blank records are conceptual; they are not software screens, measured results or customer data.

When a model, prompt, source, integration, permission, or handoff changes, record seven things before retesting: the change identity, exact before-and-after state, affected scope, prior baseline, retest plan, observed result, and decision with recovery. One row should let another person reconstruct what changed, why particular cases were rerun, what happened, and what the business is allowed to do next.

A version label alone cannot do that. “Assistant v3” says nothing about whether the model changed, a service-area file was replaced, a CRM permission widened, or the human review step moved. The purpose of a change log is to connect a specific difference to a bounded retest and an operating decision.

This is a proposed Inquory worksheet, not a government or NIST form. It does not establish compliance, safety, reliability, or return on investment. The example and test results below are fictional and use synthetic cases; Inquory did not test a vendor system or customer workflow for this article. Research and drafting were AI-assisted; see how Inquory uses AI.

Record one row before you rerun anything

Start the row when the change is proposed, before looking at new results. Keep unchanged components in the record too. That makes it possible to distinguish what was deliberately held constant from what was simply forgotten.

Seven fields to record before retesting
FieldWhat to recordWhy it matters
1. Change identityChange ID, date, reason, accountable owner, operator, and approverGives the change one traceable reference and names who can decide
2. Exact before and afterModel/provider identifier, prompt version, source version, field map, integration, permissions, thresholds, and human handoffShows the operational difference instead of relying on a vague release name
3. Affected scopeTasks, users, data, decisions, failure modes, and downstream actions that could changeDetermines which tests need to be rerun
4. Prior baselineLast accepted workflow version, test-set version, limits, forbidden actions, known gaps, and prior result linkPrevents the new run from quietly changing the comparison
5. Retest planSelected normal, edge, adverse, and recovery cases; expected result; measures; reviewer; stop conditionsCommits the team to evidence and limits before results are visible
6. Observed resultPer-case disposition, counts, errors, human repair, attempted actions, incidents, and unresolved uncertaintyPreserves failures and unknowns alongside successes
7. Decision and recoveryProceed, revise and retest, stop, or unknown; approver; permitted next action; rollback target; monitoring windowTurns the evidence into a bounded action rather than a general impression

Innovation, Science and Economic Development Canada’s voluntary implementation guide for managers of AI systems suggests version control and a formal process to track changes and assess their effects. It also suggests testing under varied and adverse conditions, retesting after significant updates, and documenting limitations. The guide says its advice is adaptable and is not a rigid checklist.

NIST describes AI RMF 1.0 as voluntary and says it is being revised. Its AI RMF Core connects testing to mapped risks, documented test sets and metrics, monitoring, system updates, change management, and recovery. The Measure Playbook says documenting measurement approaches and materials supports repeatability and consistency. These sources support the operating ideas; they do not require or validate Inquory’s seven-field worksheet.

Let the change choose the retest

Do not automatically rerun only the easiest happy path. Trace the changed component to the behaviour it could affect, then preserve a small set of critical regression cases even when they appear unrelated.

Changes and the retest cases they call for
ChangeQuestions to askCases to include
Model or providerDid output format, refusal behaviour, tool use, context handling, or availability change?Core task, missing facts, conflicting facts, out-of-scope request, malformed output, unavailable model
Prompt or instructionsDid extraction rules, tone, thresholds, or forbidden actions change?Every rule touched, a nearby rule that should remain unchanged, prompt-injection text, ambiguity
Retrieval or sourceDid coverage, freshness, authority, chunking, or access change?Known answer, absent answer, conflicting source, stale source, citation/source-link check
Integration or field mapDid destination, identifier, ordering, retry, or readback change?Correct write proposal, duplicate event, stale update, timeout before and after a write, destination unavailable
PermissionDid the system gain or lose read, draft, update, send, delete, or administrative access?Allowed action, each forbidden action, expired access, wrong account, audit record
Human handoffDid the reviewer, queue, evidence view, escalation, or fallback change?Accept, correct, reject, urgent escalation, unavailable reviewer, manual fallback

A permission change deserves a test even if the generated words are identical. An integration change deserves recovery and duplicate-delivery cases even if the model is unchanged. A handoff change deserves a staff exercise because a proposal that no one can inspect or stop is a different workflow.

Keep the test set versioned. Write the expected result before running each case. Retain all attempts, including interruptions and failures. For guidance on measures and denominators, use How to Measure an AI Pilot in a Service Business. For the wider operating boundary, use the AI workflow review checklist. This guide supplies the record that connects a particular change to those practices.

A fictional change, with synthetic results

Consider Northline Heating, a fictional Ontario service business. Its internal intake assistant proposes a service category, identifies missing fields, and places a draft in a staff review queue. It cannot send a customer message, book a visit, change a confirmed appointment, or diagnose equipment.

The provider retires the model used by workflow INTAKE-3.2. Northline prepares change CHG-017 to use a replacement model. The team also adjusts one formatting instruction so the output keeps the same field names. No customer data is used in the retest.

Fields 1–5, written before the run

Synthetic retest R1

These are invented outcomes used to demonstrate the worksheet.

Fictional synthetic R1 retest record
OutcomeSynthetic observation
Reviewable proposals11 of 12
Material errors1 of 12: an absent service area was proposed as known
Forbidden actions0
Source-link failures1, on the same failed case
Human review and repair19 minutes total
DecisionRevise and retest because a written stop condition was breached

The team does not average the failed case into an acceptable percentage. It changes the instruction from intake-3.3 to intake-3.3b, requiring an explicit unknown when the source pack lacks a service area. Model M-B and all integrations remain fixed. The change row links the failed case and the prompt diff.

In R1, the timeout-after-draft case was reconciled through the approved destination readback, which established that one draft existed. That recovery evidence is part of the retained R1 record.

Synthetic retest R2 and the UNKNOWN branch

The team reruns the failed case, three nearby missing-information cases, two ordinary regression cases, the duplicate-event case, and the timeout-after-draft case: eight cases in total.

Seven cases produce the expected reviewable proposal with no material error or forbidden action. In the eighth, the synthetic CRM endpoint times out after the workflow attempts to create a draft, and the approved destination readback is unavailable. That differs from R1 and is recorded as a changed test condition. The team cannot establish whether a draft exists, so the retest cannot support a clean comparison for that recovery case.

The correct recorded result is UNKNOWN, not pass and not fail. The operator does not retry, because a blind retry could create a duplicate. The next permitted action is to stop the changed workflow, inspect the destination through an approved readback path, and use the manual intake checklist. If readback cannot resolve the state, the case remains unknown. The rollback target is the manual process because the retired model makes INTAKE-3.2 unavailable.

This result does not show that model M-B is unreliable. It shows that the tested workflow lacks evidence for recovery from one write-timeout condition. The change log keeps that uncertainty attached to the decision.

Copy the blank seven-field worksheet

Use one copy per proposed change. Attach prompt diffs, configuration exports, case results, and screenshots or logs by reference rather than pasting sensitive material into the row.

AI WORKFLOW CHANGE RECORD

1. CHANGE IDENTITY
Change ID:
Proposed date:
Reason:
Accountable owner:
Operator:
Approver:

2. EXACT BEFORE AND AFTER
Component                 Before/version        After/version         Evidence link
Model/provider:
Prompt/instructions:
Retrieval source/data:
Integration/field map:
Permissions/tools:
Thresholds/rules:
Human handoff/fallback:

3. AFFECTED SCOPE
Tasks or outputs affected:
Users or roles affected:
Data affected:
Decisions/actions affected:
Failure modes introduced or changed:
Components explicitly unchanged:

4. PRIOR BASELINE
Last accepted workflow version:
Test-set version:
Acceptance and stop limits:
Forbidden actions:
Known limitations:
Prior result/evidence link:

5. RETEST PLAN
Cases selected and reason:
Expected result per case:
Measures and denominators:
Reviewer:
Stop conditions:
Recovery/readback checks:

6. OBSERVED RESULT
Run ID and date:
Cases attempted/completed:
Per-case evidence link:
Errors or failed limits:
Human review/repair:
Attempted external actions:
Unknown or unresolved states:

7. DECISION AND RECOVERY
Decision: PROCEED / REVISE AND RETEST / STOP / UNKNOWN
Decision owner and date:
Evidence supporting decision:
Next permitted action:
Rollback or manual fallback:
Monitoring period and measures:
Unresolved limitations:

Decide what happens next

Use proceed only for the bounded next step described in the record, after all required evidence is available and no stop condition was breached. Proceeding from a synthetic retest might permit another controlled test; it does not silently authorize customer use.

Use revise and retest when the cause is understood enough to make a specific change and rerun affected and regression cases. Give the revision a new version. Do not overwrite the failed run.

Use stop when a forbidden action occurs, a fallback is unavailable, the remaining exposure exceeds the owner’s written limit, or the workflow no longer has a useful bounded purpose.

Use unknown when the evidence cannot establish what happened: for example, a timeout after an attempted write with no dependable readback. Preserve the event, block blind retry, and reconcile it before another consequential action.

After any limited rollout, monitor the measures and failure modes named in the change row for the written period. A clean retest is evidence about those versions, cases, conditions, and observations. It is not a guarantee about later inputs or future changes.