RPO and RTO definition
Disaster recovery (DR) is the set of policies, tools and procedures for restoring systems and data after a major disruption, such as a cloud region outage, ransomware or accidental deletion. It is defined by two targets: the recovery time objective (RTO), how quickly service must return, and the recovery point objective (RPO), how much data loss is acceptable.
RPO vs RTO
The recovery point objective is measured backward from the incident: an RPO of one hour means you accept losing up to an hour of data, so backups or replication must capture changes at least that often. The recovery time objective is measured forward: an RTO of four hours means the service must be running again within four hours of the disaster being declared.
Both should come from business impact, not technology preferences. A marketing site might accept a day of RTO and RPO, while an order system may need minutes. Lower targets cost more, so many companies set tiers for different systems. Disaster recovery complements high availability, which handles smaller, routine failures automatically.
Disaster recovery strategies
Cloud providers describe a ladder of strategies, each faster and more expensive than the last. Most organizations mix them, matching each system to the cheapest strategy that still meets its agreed RTO and RPO, and documenting who may declare a disaster and trigger the switch.
Whichever strategy you choose, recovery depends on more than data. Secrets, DNS, certificates, container images, third-party allowlists and access for the on-call team all need to exist in the recovery region, and they are the items most often missed in a first drill. The four common strategies are:
- Backup and restore: regular backups copied to another region; cheapest, with recovery measured in hours
- Pilot light: core data replicated and minimal infrastructure ready, scaled up during a disaster
- Warm standby: a scaled-down but running copy of the full environment that can take traffic quickly
- Multi-site active-active: full capacity in two or more regions serving traffic, with near-zero RTO and RPO
Backups, ransomware and the 3-2-1 rule
Replication alone is not disaster recovery. If someone deletes a table or ransomware encrypts data, replication copies the damage everywhere within seconds. Point-in-time backups that are immutable and stored in a separate account are what let you rewind to before the incident. The classic 3-2-1 rule still applies: three copies of data, on two different types of storage, with one copy offsite.
Protect backups as carefully as production: separate credentials, object lock or immutable vaults such as AWS Backup Vault Lock, and alerts on deletion attempts. Ransomware groups often target backups first, precisely because backups are what make paying a ransom unnecessary.
Testing and documenting a DR plan
An untested plan is a guess. Schedule restore tests and failover drills, measure the actual recovery time against the RTO, and fix whatever slowed the team down: missing runbooks, expired credentials, DNS TTLs that were too long or dependencies nobody remembered. Infrastructure as code makes recovery far more reliable, because the environment can be recreated in another region from version-controlled templates.
Nexzem helps clients set RTO and RPO per system, implements backups and cross-region recovery on AWS, Azure and Google Cloud, and runs recovery drills, so the first full restore does not happen during a real emergency with customers waiting.