Backup and Disaster Recovery
Backup and disaster recovery work produces a recovery plan that has actually been restored from, and dated. An untested backup is a belief, and beliefs do not restore databases.
A plan you have actually executed
Almost every organisation has backups. A much smaller number have restored from them recently, and a smaller number still have restored under the conditions a real incident would impose — the primary environment unavailable, the usual administrator unreachable, and the credentials for the backup system stored in the system that is down.
So the work divides into two halves. The first is unremarkable: agree RPO and RTO per system in business terms, implement backups that meet them, hold immutable or offline copies. The second is what distinguishes it: restore, on a schedule, and record the date and the result.
Where a test fails, that is the finding. It gets written down rather than quietly repeated until it passes, because the pattern of failures is more informative than the eventual success.
What the engagement covers
Six workstreams. The fourth is the one that makes the rest real.
RPO and RTO per system
How much data loss and how much downtime is acceptable, agreed in business terms with the people who bear the consequence rather than assumed by IT.
Backup implementation
Coverage across systems, databases and configuration, with immutable or offline copies so that an attacker with administrative access cannot delete the recovery position.
Credential separation
Backup system credentials held separately from the primary identity domain, because a recovery plan that requires logging into the compromised domain is not a recovery plan.
Restore testing
Restores executed on a schedule agreed with you, each one dated with its result recorded — including the failures, which are the findings.
Recovery runbooks
Documented procedures per scenario, written for whoever is available at the time rather than for the person who designed the system.
DR rehearsal
A failover exercise at least annually, run against the assumption that the primary environment and its usual administrator are both unavailable.
How the engagement runs
Agree targets
RPO and RTO defined per system with business owners, so that recovery expectations are stated rather than inferred after an incident.
Assess coverage
Current backups checked against those targets, with gaps quantified — including systems nobody realised were unprotected.
Implement
Coverage completed, immutable or offline copies established, and backup credentials separated from the primary domain.
Test restores
Restores performed and dated, failures recorded as findings, and runbooks corrected wherever the test diverged from the document.
Rehearse and review
An annual failover rehearsal under realistic constraints, with findings written down and the plan updated.
Recovery positions compared
What each protects against, and what it does not.
| Position | Protects against | Does not protect against | Typical RTO |
|---|---|---|---|
| On-site backup only | Hardware failure, deletion | Site loss, ransomware | Hours |
| Off-site copy | Site loss | Ransomware with admin access | Hours to a day |
| Immutable or offline copy | Ransomware, malicious deletion | Slow-burn corruption | Hours to a day |
| Warm standby | Site loss, extended outage | Data corruption replicated across | Under an hour |
| Tested plan with runbooks | Human error under pressure | Nothing on its own | Depends on the above |
What we will not sign off
We will not describe a backup position as a recovery plan until a restore has been performed. Configuration that reports success is not evidence; a dated restore is. Where a client wants the documentation without the testing, we will produce the documentation and state clearly on it that the plan is unverified.
Replication also is not backup. Synchronous replication protects against losing a site and faithfully replicates corruption and malicious deletion to the second site within seconds. Where a business case rests on replication alone, that gap needs stating before an incident rather than after.
- An untested backup is a belief. A dated restore is evidence.
- Hold backup credentials outside the primary identity domain.
- Replication protects against site loss, not against corruption or ransomware.
- A failed restore test is the finding; record it rather than quietly retrying.
Questions buyers ask about this
How often do you test restores?
On a schedule agreed with you, typically quarterly for critical systems and annually for the rest. Every test is dated with its result recorded. If a test fails, that is the finding — it is written down rather than repeated quietly until it passes, because the failure pattern tells you more than the eventual success.
What about ransomware specifically?
Immutable or offline copies so that administrative access cannot delete the recovery position, credential separation so the backup system is not reachable from the compromised domain, and a recovery rehearsal that assumes the primary identity domain is unavailable. Most recovery plans fail that last assumption.
Is replication the same as backup?
No. Replication protects against losing a site and will faithfully copy corruption or malicious deletion to the second site within seconds. Both have a place; treating replication as a backup is one of the more common and more expensive misunderstandings in this area.
How do you set RPO and RTO?
With the business owners of each system, in terms of acceptable data loss and acceptable downtime, because those are commercial judgements rather than technical ones. IT can state what each target costs to achieve; only the business can say what the outage costs.
What does a DR rehearsal involve?
A failover exercise run under realistic constraints — the primary environment unavailable and the usual administrator not participating — with the runbook followed as written. Every point where the runbook proves inadequate is a finding, and correcting those is the actual output of the exercise.
Marked up as FAQPage structured data, matching the visible text exactly.
Still not sure this is the right service?
Answer four questions and we will tell you which one fits — or that none of them do.
Tell us what you are trying to fix.
A short conversation about the objective, the constraints and the timing. If we are not the right fit, we will say so.