
AI and Automation
Who Is Responsible When an AI Agent Takes Action?
An agent can take action. Someone still has to answer for it.
Read the perspectiveWhat changes when AI can act
A writing assistant produces a draft that someone can review. An agent connected to business tools may search a customer record, interpret a request, choose a next step and update a system. The distance between a suggestion and its consequences becomes shorter.
That difference matters more than the label used in a product demonstration. Ask what the system can actually do. Can it read information only? Can it prepare an action for approval? Can it execute that action? Can it choose a sequence of further actions after seeing the result?
Anthropic distinguishes agents from predefined workflows by the model’s role in directing the process and using tools. A business evaluating these systems should turn that architectural distinction into an operating question: how much discretion is being delegated, and over which work?
The answer may vary within one process. A support agent might be allowed to retrieve delivery status and draft an explanation while remaining unable to change an address or issue a refund. That limited role can still be useful. An agent does not need broad authority to justify its existence.
Responsibility must remain clear even when the work crosses several systems.
Follow a request through the shared workflow
Suppose a small online retailer is considering an agent to handle order-status requests. This is an illustrative design, not a report of an MT BYTES deployment.
The customer asks why a parcel has not arrived. The agent identifies the relevant order, retrieves tracking information and prepares a response. If the carrier record is consistent and the order is within the stated delivery window, the business might permit a routine update. If the customer disputes receipt or asks for compensation, the request moves to a person with the relevant authority.
Several responsibilities sit behind that apparently simple exchange. Someone decides which records the agent may access. Someone maintains the delivery policy. Someone determines when a case must be escalated. Someone checks whether the resulting messages are accurate and appropriate. Someone responds when the connection to the carrier fails.
These jobs do not disappear because the customer sees one conversation. They become part of the service’s design and operation.
A useful responsibility map names roles rather than departments. In a small business, one person may hold more than one role, but each responsibility still needs to be explicit:
| Responsibility | Practical question |
|---|---|
| Business owner | What outcome is this service meant to improve? |
| Policy owner | Which rules may the agent apply, and who changes them? |
| Authorised approver | Who can approve actions outside the routine boundary? |
| Technical operator | Who maintains connections, permissions and monitoring? |
| Quality reviewer | Who examines outputs, exceptions and complaints? |
If a row has no owner, the agent is being asked to operate inside an organisational gap.
Treat tool access as delegated business authority
Connecting an agent to a system is not simply a technical installation. Access determines which information it can see and which consequences it can create.
Start with the minimum actions needed for the agreed task. Reading an order does not require the ability to delete it. Preparing a refund does not necessarily require issuing one. Access to one customer record does not justify unrestricted access to an entire customer database.
Set boundaries in the tools and surrounding software where possible. A written instruction to “be careful” is weaker than an operation that refuses an unauthorised action. Define permitted fields, transaction limits, customer scope and approval requirements as enforceable rules.
The surrounding environment matters alongside the model. Anthropic’s account of trustworthy-agent design considers the model, operating framework, tools and environment together. That is a useful corrective to evaluating an agent only by how convincing its answers sound.
Also examine indirect access. An agent may encounter instructions inside emails, documents or webpages supplied by other people. Those materials should be treated as task data, not as authority to change permissions or business policy. Test how the system behaves when information is conflicting, incomplete or deliberately misleading.
Record consequential actions with enough context for a reviewer to reconstruct what happened. Include the request, relevant evidence, the action taken and any approval. Logging should be proportionate and protect sensitive information rather than become a second uncontrolled copy of customer data.
Give people a workable escalation role
“A human will review it” sounds reassuring until the review queue grows faster than anyone can manage. Human oversight needs capacity, useful context and a clear decision.
An approver should see what the agent proposes, why the case requires review and which records support it. The interface should make disagreement practical. If rejecting a proposal requires reconstructing the entire case, approval can become the path of least resistance.
Define what happens while the reviewer is unavailable. Some work can wait. Some requires an alternative route. A request should not silently become an automated approval merely because a deadline has passed unless the business has deliberately authorised that behaviour.
Set the agent’s escalation rules around meaningful uncertainty and consequence. It may need to hand over when sources disagree, an identity check fails, a request falls outside policy or a proposed action is difficult to reverse. These triggers should be tested with real task variations.
NIST’s AI Risk Management Framework places roles and responsibilities within AI governance. For an SME, this can be a short, maintained operating document rather than a new committee. The essential requirement is that people know what they own and can act on it.
Responsibility must remain clear even when the work crosses several systems.
Measure the service, including the supervision
The business case should include the work transferred to reviewers and operators. If an agent handles routine requests quickly but produces ambiguous escalations, the remaining workload may become harder. Counting completed automated actions alone will not reveal that change.
Track whether the intended task is completed correctly, how often people intervene, which exceptions recur and what it takes to correct an error. Examine the experience of the customer and the employee receiving the handoff. A fast first reply has limited value if it creates another unresolved exchange.
Prepare employees for the changed division of work. Show them the agent's authorised scope, the situations it will hand over and how to report a questionable action. The aim is informed use: neither accepting every output automatically nor repeating every completed task because nobody knows what the system has checked.
Review permissions and policies when the service changes. Adding a new tool, product range or market may extend the consequences of an existing action. A permission that was reasonable for a narrow pilot can become inappropriate when the agent begins serving more customers.
Begin with a bounded service and expand authority only when evidence supports the next step. Retain a way to stop automated actions while preserving access to the information needed to continue the work manually.
An AI and automation engagement should therefore define both the technical system and the operating responsibilities around it. The practical outcome is a service whose authority is understood, whose exceptions are visible and whose owner can explain how it is performing.
