Database Replication definition
Database replication is the process of continuously copying data from one database server, the primary, to one or more others, called replicas, so the same data exists in several places. Replication improves availability, because a replica can take over if the primary fails, and scalability, because read queries can be spread across replicas.
How database replication works
Most databases record every change in a log before applying it: the write-ahead log in PostgreSQL, the binary log in MySQL, the oplog in MongoDB. Replication ships that log, or the changes it describes, to replica servers, which replay it to stay in step with the primary. Applications write to the primary and can read from any replica, usually through separate connection strings or a proxy that routes reads.
Managed services hide much of the work. Amazon RDS and Aurora, Azure Database, Google Cloud SQL and MongoDB Atlas let teams add read replicas or cross-region copies with a few clicks, and promote a replica to primary during failover. The underlying trade-offs, however, remain the same and still shape application design.
Synchronous vs asynchronous replication
With synchronous replication, the primary waits until at least one replica confirms a write before reporting success. No acknowledged data is lost if the primary fails, but every write pays an extra round trip, so synchronous replicas usually sit in the same region. With asynchronous replication, the primary confirms immediately and replicas catch up shortly after, which is faster and works across long distances, but recent writes can be lost in a failover.
Many setups mix both: a synchronous standby in another availability zone for high availability, plus asynchronous replicas in other regions for reads and disaster recovery. The choice directly expresses how much data loss, measured as RPO, the business can accept.
Replication topologies
Replication can be arranged in several ways, each with its own balance of simplicity, write capacity and risk of conflicting changes. Most applications need only the first pattern for years, and should add complexity only when measurements demand it:
- Primary-replica (single leader): one writable primary and read-only replicas; the most common and simplest pattern
- Multi-primary (multi-leader): several nodes accept writes, useful across regions, but concurrent edits can conflict and need resolution rules
- Leaderless: any node accepts writes, and quorum reads and writes keep data consistent, as in Cassandra
- Cascading: replicas feed other replicas, reducing load on the primary in large fleets
Replication lag and its pitfalls
Asynchronous replicas are always slightly behind, normally by milliseconds but sometimes by seconds or minutes under heavy load or long-running queries. A user who updates a profile and is immediately shown a page read from a lagging replica may see the old value and assume the save failed. Fixes include reading your own writes from the primary for a short period, routing by session, or checking replica position before reading.
Replication is also not a backup. A mistaken DELETE or corrupted data replicates to every copy within moments, so point-in-time backups are still required. Teams choosing a database for distributed workloads should understand how replication interacts with the CAP theorem. Nexzem configures replication, failover and lag monitoring as part of database and cloud architecture work.