
Cloud Engineering
How Much Control Should AI Have Over Cloud Operations?
Set clear limits before AI changes a live cloud environment.
Read the perspectiveBreak the operating task into distinct steps
During an incident, a cloud operations team may need to collect logs, compare recent deployments, inspect dependency health and decide what to change. These activities are related, but they carry different consequences.
Gathering evidence can be time-consuming. AI assistance may help assemble relevant information and suggest where to investigate. Executing a change introduces another level of authority, because the action can affect customers, data and costs.
AWS’s DevOps Agent documentation describes correlating operational information such as telemetry, code and deployment context. Evaluate such capabilities against a defined task, rather than treating the product category as proof that unattended operation is ready.
Separate the proposed service into observation, explanation, recommendation and execution. Ask what evidence supports each stage and which permissions it requires.
An SME might begin with an assistant that prepares an incident brief for an operator. That can be useful without giving the assistant authority to restart services or change access. The next step should follow evidence from use and a deliberate decision about additional responsibility.
The operating objective is faster, better-informed action with controlled consequences. Autonomy is a design choice within that objective.
Every authorised automated change needs a defined way to verify its effect and stop further action.
Choose a task an operator can assess
Start with a task whose output an operator can assess. Summarising relevant alerts, identifying recent changes or preparing a comparison of error patterns may offer a manageable boundary.
Define the sources the system may inspect and the time period relevant to the investigation. Give it enough context to distinguish production from test environments and expected maintenance from unexpected failure.
Set a clear completion condition. “Investigate the incident” is broad. “Collect the affected service’s recent errors, deployment changes and dependency status, then identify unresolved questions” produces a more reviewable result.
Test the assistant against known incidents and ordinary variations. Include cases where the available evidence does not establish a cause. A useful investigation should be able to state uncertainty rather than invent a confident explanation.
Measure the operator’s work. Does the brief reduce the time spent finding information? Are the cited records relevant and current? Does it omit a critical source? How much verification is needed before the operator can use it?
Keep the result connected to the original evidence. An explanation without a practical route to inspect its supporting records may add another layer of checking rather than reduce effort.
Avoid beginning with a task chosen only because it looks impressive in a demonstration. The business needs a repeatable improvement to real operations.
Match controls to each level of authority
Read access should be scoped to the information the task needs. Operational logs can contain sensitive data or secrets, so broad access is not harmless simply because the agent cannot modify infrastructure.
Recommendations should identify the proposed action, affected resources and assumptions. The operator needs to understand why the action is relevant and what could happen if the explanation is wrong.
Execution needs stronger controls. Define which operations are allowed, under what conditions and within which resource or cost limits. Enforce those limits through the surrounding tools and permissions.
| Operating level | Example | Required boundary |
|---|---|---|
| Inspect | Read selected service logs | Scope, retention and sensitive-data handling |
| Explain | Summarise a likely dependency failure | Traceable evidence and visible uncertainty |
| Prepare | Assemble a proposed configuration change | Reviewable difference and expected effect |
| Execute | Perform a narrowly authorised action | Enforced permission, stop condition and action record |
AWS guidance on human review addresses risk-based approval and escalation. Apply that to the consequence of the change, rather than using one approval rule for every operation.
A restart, deployment rollback, permission change and database modification are not interchangeable. Each needs a specific operating decision.
Design recovery from the automated action
Before authorising execution, ask how the team would detect and correct a bad action. A successful API response does not establish that the service improved.
Define the expected effect and the checks that follow. If an action changes capacity, inspect service performance and cost. If it rolls back an application, confirm that the data and dependent services remain compatible.
Rollback itself needs careful treatment. Some changes can be reversed directly; others have consequences that require a separate recovery process. An automated database change should not be described as safe merely because a backup exists.
Set limits on repeated action. An agent should not keep restarting a service or adding capacity indefinitely because the original symptom remains. Escalate when the action fails to produce the expected result or the evidence changes.
Record the input, recommendation, approval and executed operation at a level appropriate to the service. Protect sensitive information in those records. An operator should be able to reconstruct what happened without relying on a conversational summary alone.
Every authorised automated change needs a defined way to verify its effect and stop further action.
Test the stop mechanism under realistic permissions. The person accountable for the service must be able to suspend execution when the system behaves unexpectedly.
Check the authority required to operate the stop mechanism. If suspending the agent depends on the same unavailable identity or connection that caused the incident, the fallback may fail when needed. Maintain an appropriate independent access route and review it with the other emergency operating arrangements.
Keep operational evidence separate from authority
An operations agent may inspect logs, tickets, code comments or external documents created by many people. Some content can be incomplete, misleading or malicious.
The system should not treat instructions found inside that material as permission to change its own rules or access. A log entry is evidence about the service, not an authorised command to the operations agent.
The NCSC guidance on careful adoption of agentic AI highlights risks associated with components, integrations and downstream actions. Review the whole path from information source to potential change.
Use trusted channels for policy and permission decisions. Keep the agent’s operating identity separate from individual administrator accounts where the environment supports it, and grant only the required capabilities.
Review external tool connections before adding them. A new integration can expand both the information exposed and the actions possible. The original assessment may no longer describe the service after that change.
Include adversarial and conflicting inputs in testing. The objective is to verify that the system preserves its authorised boundary when the evidence it reads contains something unexpected.
Keep an operator responsible for the service
Assign an owner who understands the supported workload, can evaluate escalations and has authority to stop automated actions. The owner also needs time to review performance and maintain the arrangement.
Track the quality of investigations, unnecessary escalations, incorrect proposals and the outcomes of executed changes. Count the supervision required. An agent that produces many plausible but weak recommendations may increase operational workload.
Re-evaluate after changes to the model, tools, permissions or infrastructure. Maintain a set of representative cases that can reveal whether important behaviour has changed.
Begin with a bounded deployment and expand only when the evidence supports the next level of authority. Retain a useful read-only mode where practical so that execution can be paused without losing all assistance.
At any point, the operator should be able to establish what the agent inspected, what it changed and how to stop its next action. Keep that ability intact as the automation expands.
