All articlesDistributed Systems · Multi-cloud

Active-Active Multi-Cloud Database Replication: Overcoming Network Partitions

Active-active replication across cloud providers buys provider-level fault isolation and local write latency. It also imports every hard problem in distributed systems at once: partitions between providers are longer and less predictable than partitions within one, and both sides remain fully operational while they last. The architecture only works when partition behaviour is a deliberate decision rather than an emergent one.

Decide what a partition should do

During a cross-provider partition, each side can continue accepting writes and reconcile later, or refuse writes to preserve a single ordering. There is no third option, and the answer differs by table. Ledger entries and reservations usually cannot tolerate divergence; session state, preferences, and telemetry usually can.

Writing this decision down per dataset, before deployment, is what separates a designed system from one that discovers its own semantics during an outage.

Make conflict resolution explicit

Default last-writer-wins is a policy, not the absence of one, and it silently discards data under clock skew.

  • Use hybrid logical clocks rather than wall-clock timestamps for ordering
  • Apply CRDT semantics to counters, sets, and additive fields
  • Partition ownership by region or tenant so most keys have one writer
  • Log every resolved conflict to an audit stream for later review

Connectivity and routing

Dedicated interconnects between providers reduce jitter substantially compared with public transit, but they do not eliminate partitions — they change their frequency and shape. Assume asymmetric failures where one side can reach the other but not the reverse, and build health checks that detect them rather than reporting success from a one-way probe.

Global routing should send users to their nearest healthy write region and drain a region cleanly when it is isolated, with drain and return rehearsed as a routine operation.

Operational discipline

Multi-cloud active-active multiplies operational surface: two sets of IAM, two observability stacks, two failure vocabularies. Unify telemetry into one view keyed by logical region rather than provider account, run partition game days on a schedule, and measure convergence time after each exercise. A topology nobody has rehearsed is a topology nobody can operate at three in the morning.

Key takeaways

  • Choose availability or consistency per dataset, in writing, before launch
  • Replace implicit last-writer-wins with explicit, auditable conflict policy
  • Design for asymmetric partitions, not just clean link failures
  • Rehearse region drain and measure convergence after every game day

Talk to KodeSync Resources

KodeSync Resources engineers distributed database synchronization, multi-master replication and high-availability data layers. Send us your environment and we will respond with a scoped audit plan.

Request system audit

Related articles