Successful job, failed recovery
Backup software reports on whether data was copied. It does not report on whether that copy can be brought back into a working system, in the right order, in an acceptable time. Those are different questions.
Agree the two numbers first
Before designing anything, agree two things in plain language.
- How much recent work could you afford to redo? (This sets the backup frequency — the RPO.)
- How long could you operate before systems must be back? (This sets the recovery design — the RTO.)
Know what is not covered
Coverage gaps are usually discovered during an incident. Write the list down while things are calm.
- Cloud data such as Microsoft 365 mail, SharePoint and OneDrive
- Line-of-business application databases
- Configuration for firewalls, switches and access points
- Local files that never leave a laptop
Test in a way that proves something
A useful test restores real data to a usable state and involves someone from the business confirming the content is correct. Record the date, what was restored and how long it took.
Keep the recovery plan short
One page: recovery order, where backups live, who has access, and who to call. If it takes longer than a few minutes to read during an incident, it will not be read.