A displaced optical impression returns to one coherent blue registration.

Security, Cloud and Resilience

Put Your Disaster Recovery Plan to the Test

A backup only helps if your team can restore the service.

MT BYTES6 min read
Read the perspective

Define what recovery means for the business

“The system is back” can mean several things. The server may be running while staff remain unable to sign in. The database may be restored while recent transactions are missing. The website may load while its payment connection fails.

Choose the business task that must resume and describe a successful recovery in those terms. An authorised employee can retrieve the current booking, make an agreed change and confirm it to the customer. A finance user can access verified records and complete the required payment process. These statements are testable.

Agree the maximum acceptable interruption and the amount of recent information that could be lost. Recovery time and recovery point objectives describe these different needs. AWS’s business-continuity guidance links them with the requirements of the business.

Avoid selecting targets solely because a supplier offers them or because a short number sounds reassuring. A very demanding target can carry substantial cost. An undemanding one may be incompatible with customer commitments. The business owner and technical team need to resolve that trade-off together.

Record the circumstances. Recovery outside normal working hours may depend on different staff and supplier support from a weekday rehearsal.

Recovery is complete when the business can resume the agreed task with information it has reason to trust.

Find the dependencies a backup leaves out

A backup protects specified information or system state. Recovery may also need credentials, encryption keys, deployment instructions, network configuration, licences and access to external services.

List those dependencies before the exercise. Establish where they are held and who can use them if the normal environment is unavailable. A document stored only inside the affected system cannot be the sole recovery instruction.

Check the relationship between components. Restoring an application and its database from different points may create inconsistencies. A message queue may contain work already reflected in another record. An external provider may have processed a transaction while the local system failed to record the response.

These details influence the order of recovery and the validation required afterwards. They should be understood by the people responsible for the service, with specialist input where necessary.

Include the human dependencies. Does one person hold essential account access? Can another authorised person perform the procedure? Are supplier contact details current? Is there a decision-maker who can authorise a temporary workaround?

A useful recovery plan connects the technical steps with the people and decisions required to complete them. It should be detailed enough to use under pressure without becoming a document nobody maintains.

Design a rehearsal with a clear boundary

Choose a scenario that exercises an important part of the plan without creating unnecessary disruption. A controlled restoration into an isolated environment may be appropriate for an initial test. More complex exercises require a defined scope, precautions and approval from the relevant owners.

State what will be unavailable in the scenario. Otherwise participants may unconsciously rely on the very system or person the plan is supposed to replace. If the exercise assumes the primary identity service is unavailable, using it to retrieve every credential leaves that assumption untested.

Define the starting event and the stopping condition. Begin the clock when the exercise starts, and stop when the agreed business task can be completed with validated information. Record intermediate times so delays can be understood.

AWS’s recovery-testing guidance emphasises testing as a source of confidence in the strategy. The practical value comes from observing the steps, not merely recording that an exercise occurred.

Assign an observer where the team’s size allows. Capture missing instructions, access problems, manual decisions and unexpected dependencies. The aim is to improve the arrangement rather than reward participants for improvising quickly.

Protect the test environment and any restored information. A recovery exercise should not create an uncontrolled copy of sensitive production data.

Validate the work after restoration

Technical restoration is followed by business validation. Check that the relevant records are present, internally consistent and usable through the application.

For an illustrative appointment service, the checks might include retrieving a known booking, confirming its customer and time, making a permitted change and checking that the change appears in the right place. The business should also examine how it will handle appointments received while the system was unavailable.

Choose representative checks before the exercise. Include a normal transaction, an exception and a recently changed record. Where financial or sensitive information is involved, the validation needs to reflect those consequences.

Record what information is missing and how its absence will be handled. A recovery point objective describes tolerated loss; it does not make the missing work disappear. Someone must decide how to reconstruct, confirm or communicate affected transactions.

Check integrations carefully before resuming automated actions. A restored service may attempt to send messages or repeat jobs that have already happened. The restart procedure should avoid turning a successful restore into duplicate operational work.

Recovery is complete when the business can resume the agreed task with information it has reason to trust.

Only then can the team compare the observed result with its target and decide what needs to change.

Rehearse communication and temporary operations

The business may need to operate while recovery continues. Decide which activities can proceed manually, which must pause and who can authorise exceptions.

A workaround should specify the information used, the record kept and the method for reconciling later. A spreadsheet created during an outage can become another source of conflicting data if no one owns the return to normal operation.

Communication needs similar clarity. Identify who informs staff, suppliers and customers, what can be said with confidence and when the next update will occur. Avoid promising a restoration time merely because someone wants an answer.

Separate internal technical updates from customer information. Customers usually need to know the effect on their service, what action they should take and how the business will keep them informed. They do not need speculative technical explanations.

Regulatory, contractual and incident-reporting obligations must be established for the relevant business and jurisdictions. Keep the required contacts and advice route available before an incident, rather than attempting to resolve every obligation during the interruption.

Microsoft’s reliability-target guidance connects service objectives with business expectations and measurement. Communication commitments should be consistent with those operating capabilities.

Turn rehearsal findings into recovery improvements

A failed rehearsal can be valuable if its findings lead to completed work. List each gap, its effect, the responsible owner and the next check.

Distinguish problems in the procedure from problems in the underlying design. Missing instructions may be corrected quickly. An unavailable key, an incompatible backup or an unrealistic dependency on one supplier may require a larger change.

Retest the affected part after the repair. Keep the previous result so the business can see whether the gap closed. Do not replace an observed failure with a statement that it should now work.

Review the plan when the application, infrastructure, supplier arrangement or staffing changes. A procedure that worked for an earlier architecture may no longer match the service.

A cloud engineering engagement can help turn a recovery requirement into a testable technical and operating arrangement. The deliverable should include the evidence from the rehearsal and the remaining work, giving the business a practical view of its ability to resume service.

MT
MT BYTES

Perspectives on technology and business.

Explore perspectives

Find out whether your recovery plan can restore the service

MT BYTES can help scope a recovery review or controlled rehearsal around a critical application. Bring the business task, the current backup arrangements and the interruption your team needs to prepare for.

Discuss your project