All articlesDistributed Systems · PostgreSQL

Optimizing PostgreSQL Replication Lag in High-Write Environments

Replication lag in PostgreSQL is rarely one problem. It is three, and they need to be told apart before anything is tuned: the primary generating more WAL than it can ship, the network failing to move it, or the replica failing to apply it fast enough. Treating apply lag with network tuning wastes a maintenance window and leaves the cause untouched.

Separate write, flush, and replay lag

pg_stat_replication exposes write, flush, and replay LSNs per standby. Comparing each against the primary's current LSN localises the bottleneck: a gap at write points to shipping, a gap between flush and replay points to the replica's recovery process.

Because replay is single-threaded in PostgreSQL's physical replication, a replica can keep up with network delivery and still fall steadily behind on apply — the most common pattern in high-write environments.

Reduce WAL at the source

The cheapest lag improvement is generating less WAL in the first place.

  • Tune checkpoint frequency to limit full-page writes after checkpoints
  • Drop unused indexes — every index multiplies write volume
  • Batch small transactions and avoid update-heavy churn on wide rows
  • Enable WAL compression where CPU headroom allows

Help the replica keep up

Replicas need I/O capacity comparable to the primary; provisioning them on cheaper storage guarantees apply lag under load. Enable recovery prefetching so the startup process issues reads ahead of replay, and watch for query conflicts: long-running analytics on a hot standby cause recovery pauses unless feedback and delay settings are configured deliberately.

Where a single replay thread is genuinely the ceiling, logical replication with multiple subscriptions can parallelise apply across tables, at the cost of losing full-cluster physical fidelity.

Decide what lag means to the application

Some lag is acceptable; unbounded lag is not. Define a tolerance, route reads accordingly, and fail read traffic away from a standby that exceeds it rather than serving stale results silently. For read-your-writes correctness, pin a session to the primary or gate reads on an observed LSN — application-level correctness cannot be delegated to replication tuning.

Key takeaways

  • Distinguish write, flush, and replay lag before tuning anything
  • Cut WAL volume at the source — indexes and checkpoints dominate
  • Give replicas primary-class I/O and enable recovery prefetching
  • Route reads by a defined lag tolerance instead of hoping for freshness

Talk to KodeSync Resources

KodeSync Resources engineers distributed database synchronization, multi-master replication and high-availability data layers. Send us your environment and we will respond with a scoped audit plan.

Request system audit

Related articles