Illustrative report: the summary loses a promised callback
An invented support team uses AI to prepare a summary after an enquiry. A staff member notices that the summary contains the customer question but omits a callback that a colleague promised. The staff member has been adding that detail manually. This example is a process design exercise, not an account of a Yes AI customer incident.
A useful report reads: case reference EXAMPLE-17; the source message says the customer should receive a callback after the technician checks availability; the summary records the fault but has no next action; the operator had to reopen the source and add the callback manually. The report points to the authorised source record rather than copying the whole customer conversation into an open channel.
The reviewer checks whether the callback was actually promised, whether it was conditional and whether the instruction was still current. This matters because a fix that adds a callback whenever the word appears could invent work that was only discussed as a possibility. The desired result is a summary that accurately distinguishes a commitment, a conditional next step and an unanswered request.
The team preserves the original example and adds neighbouring cases: a confirmed callback, a callback dependent on approval, a customer asking for a callback without agreement, and a message with no action at all. The proposed change must improve the original summary without turning those other cases into false commitments. The test checks the meaning of the output, not whether it contains a particular phrase.
After the change is released, the owner verifies the behaviour and tells the reporter that callback commitments are now represented in a separate next-action field. The owner also checks whether any earlier cases still need a callback. Fixing future summaries and completing already-promised work are recorded as separate actions, so the technical release does not conceal an unresolved customer task.
At the next operations review, the team looks for repeat reports in the same category and checks the effort staff spend correcting summaries. If the problem continues, the scope or review arrangement may need to change. The useful result is a better operating process supported by examples, not simply a higher number of tickets marked closed.