
AI and Automation
What Happens When AI Automation Cannot Finish the Job?
Design the handoff for the cases automation cannot safely resolve.
Read the perspectiveThe happy path can leave work unfinished
Picture a hypothetical supplier receiving a request to change a delivery address. The message looks ordinary, so an automated assistant identifies the order and attempts the update. The warehouse system refuses because dispatch has begun. The assistant has understood the request correctly. The business still has an unresolved delivery problem.
What happens next matters more than the initial classification. The customer may need a clear explanation, the warehouse may need an urgent call, and the original address must remain intact until a valid alternative is agreed. Marking the automation “failed” accomplishes none of that.
A weak exception path tends to create one of two outcomes. Work stops without a visible owner, or the system continues by improvising around missing authority. The first leaves customers waiting. The second can create commitments the organisation cannot honour.
An exception is still a customer request, and somebody still owes it an outcome. Treat that sentence as a design requirement. Every branch of the process should lead to completion, a meaningful request for clarification, an assigned human task or an explicitly recorded cancellation. A technical error message belongs inside that operational story, not at its end.
An exception is still a customer request, and somebody still owes it an outcome.
Different failures need different responses
An ambiguous message, a missing record and a failed write to another application are different conditions. Sending all three to a generic “AI could not process this” queue discards information the team needs.
Ambiguity calls for clarification. If a customer asks to amend “the last order” and several records fit, the system should seek a discriminating detail. Missing information may require retrieval or a request to the customer. An unavailable application may justify a controlled retry. A policy conflict needs someone with the authority to resolve it.
There is also the problem of a convincing answer built on insufficient evidence. NIST's Generative AI Profile identifies confabulation as a risk. In operational terms, fluent language cannot substitute for a verified order status, price or permission.
Design the response around what is known. If the order exists but dispatch status is unavailable, preserve the confirmed identity and make the missing status explicit. Do not force the next person to repeat the entire investigation.
This classification can remain simple. A small business may begin with a handful of exception reasons that staff can distinguish reliably. Add detail when it changes the action someone takes. Elaborate labels with identical handling only make the queue harder to maintain.
A handoff is a package of work
A notification saying “please review” gives a colleague responsibility without context. A workable handoff should state the customer's intended outcome, the information already verified, the action attempted and the point where progress stopped. Include the relevant record and the reason human involvement is needed.
For the address-change example, a reviewer should see that the order was identified, dispatch prevented the change and no replacement address was saved. They should also know whether the customer has already received an acknowledgement. Without that detail, a well-meaning colleague may repeat an action or send a contradictory message.
The queue needs an owner and an expectation for response. Routing everything to the founder might work during a small trial, but it makes the founder the capacity limit as usage grows. Assign by the authority required: operational staff for missing details, a manager for an exception to policy, technical support for a broken connection.
A handoff also needs a return path. After review, does the person complete the request manually, approve a proposed action or send it back with corrected information? Each route should leave a record of the final outcome.
AI automation becomes easier to operate when review is designed as ordinary work with visible status, rather than an emergency escape hatch hidden from the main process.
Check what happened before repeating an action
A timeout does not necessarily mean an action failed. An application may save a booking and then fail to return confirmation. If the automation simply repeats the request, it could create another booking. The same concern applies to charges, emails, fulfilment instructions and account changes.
The recovery design should answer a precise question: can the system establish whether the first attempt took effect? Where supported, use a stable request identifier and check the resulting record before trying again. Where that is impossible, put the case into a state that calls for verification.
Separate actions that are safe to repeat from actions that need additional control. Retrieving the latest status is usually different from issuing a refund. A delayed response should not grant permission to perform a consequential action again.
Recovery may also require compensating work. If one application was updated and another was not, somebody needs to reconcile the mismatch. Keep enough history to identify both sides and determine the correct final state.
Anthropic's discussion of trustworthy agents describes controls around an agent's authority. For business implementation, permissions should also reflect recovery needs: a system allowed to create something does not automatically need unrestricted power to delete, refund or override it.
Make this behaviour visible to customer-facing staff. If a payment or booking is awaiting verification, they need language that accurately describes that state. A premature success message can turn a recoverable technical delay into a broken promise.
Test with the cases staff remember
Routine examples establish that a process can work. Difficult examples establish whether it can be operated. Ask experienced staff for anonymised cases involving conflicting instructions, incomplete records, changed circumstances and repeated messages.
Include failures in connected systems. Test a request received twice, a response arriving late, a record edited by a colleague during processing and a tool returning an unexpected format. These tests reveal whether the automation preserves the customer's intent when its assumptions stop holding.
Measure more than classification accuracy. Count how many exceptions arrive with sufficient context, how long they remain unassigned, how often staff have to reconstruct prior actions and how many cases reopen after apparent completion. A system can generate excellent summaries while maintaining a poor queue.
Architecture affects this work. Anthropic distinguishes predefined workflows from agents that choose their own steps. Greater freedom can be valuable, but it also creates more possible paths to inspect. Use flexible behaviour where the task requires it, and explicit rules where a stable process already exists.
Avoid treating a model's confidence score as permission to act. Test the proposed control against actual cases, especially those with serious consequences. A confident error remains an error, and a numerical threshold does not explain who will deal with its effects.
Read the exception queue as business evidence
Once the automation is running, review recurring exceptions with the people who resolve them. Repeated clarification requests may point to a poor intake form. Conflicting answers may expose outdated policy documents. Frequent access failures may indicate an integration or account-management problem.
Some exceptions should lead to a change in the automated path. Others should remain deliberately human. A valuable customer asking for an unusual commercial arrangement may need judgement and a conversation, regardless of how capable the system becomes.
Track the cost of maintaining the exception path when judging the project. Review time, reconciliation and customer follow-up belong in the same account as the time saved on routine cases. Otherwise the automation can appear efficient simply because its unfinished work has moved elsewhere.
The operational aim is a process that remains understandable when it encounters difficulty. A colleague should be able to see what happened, what remains unresolved and which action will move the request forward. That is the standard to meet before increasing volume or authority.
