Worked example: an invoice workflow with a misleading success rate
Imagine a fictional business reviewing its invoice intake automation. The dashboard reports that nearly all processed documents completed successfully. The accounts operator says the system is useful, but also describes a separate spreadsheet used to track invoices that never appeared in the workflow. The review owner recognises that the dashboard denominator includes only documents the automation captured, so it cannot establish intake completeness.
The team defines an invoice obligation as the business unit and compares the authorised source inbox inventory with workflow records and destination entries. Resent attachments are grouped for review rather than automatically counted as new obligations. Documents still awaiting a business decision remain held. Items with uncertain destination outcomes remain unknown until inspected. The resulting report describes the whole selected process boundary, including the work previously absent from the dashboard.
Next, the operator records review and correction time for a defined observation sample. The team distinguishes a clean automatic result from an entry corrected before posting. It discovers that one supplier format requires repeated manual account selection. The improvement proposal is therefore a specific mapping and review-rule change, with representative test invoices, rather than a broad promise to increase the model's intelligence.
The owner chooses two decisions for the next period: establish an intake reconciliation that detects missing obligations, and test the supplier mapping change before allowing it to affect posting. Each has a named owner and a completion test. The model bill is retained as one cost line, but the business case also includes staff review and correction effort. No cash saving is claimed merely because some keystrokes disappeared.
At the next review, the team checks whether the reconciliation detected the deliberately controlled missing-intake case and whether ordinary source records were accounted for. It inspects the supplier test outcomes and a bounded set of subsequent results. The meeting can now decide whether to retain, revise or expand the change using observable evidence rather than a dashboard percentage whose population was incomplete.
A useful review packet for this example contains a short decision page, the reconciled counts, the selection method for reviewed invoices, the actual examples and the current change register. The supporting material remains available to authorised staff, but the meeting does not require every participant to read every log line. Summaries point to source evidence rather than replacing it. If a number cannot be reconstructed from the supporting records, it is labelled provisional and kept out of a stronger claim.
The owner also records what the review did not establish. The sample may not represent every supplier or seasonal workload, and a controlled missing-intake probe does not prove uninterrupted coverage for an entire month. Those limitations guide the next measurement. They do not make the exercise useless; they stop a narrow observation from turning into an unsupported guarantee and keep the next decision proportionate to the evidence.