All articlesDistributed Systems · Failure modes

Solving Split-Brain Scenarios in High-Availability Distributed Clusters

Split-brain is the failure where two halves of a cluster each believe they hold leadership and both accept writes. The cluster stays up, dashboards stay green, and divergence accumulates silently until reconciliation becomes a manual, lossy exercise. It is the most expensive failure in high-availability database operations precisely because nothing appears broken while it happens.

Quorum arithmetic is the first defence

A cluster can only avoid split-brain if at most one partition can assemble a majority. That requires an odd number of voting members and a topology where no single network fault produces two equal halves. Even-numbered clusters split cleanly down the middle, which is exactly the condition to avoid.

Two-site deployments deserve particular scepticism: whatever the node count, a link failure between sites leaves two groups, and only a third location can break the tie.

Witness placement decides real behaviour

A witness or arbiter carries a vote without carrying data, and its location determines which failures the cluster survives.

  • Place the witness in a third failure domain, never alongside a data node
  • Verify the witness has independent network paths to both sites
  • Monitor witness reachability as a first-class alert, not a secondary metric
  • Test the cluster's behaviour when the witness alone is unreachable

Fencing turns detection into safety

Quorum tells a node it should not be leader; fencing makes it impossible for it to act as one. STONITH-style power fencing, storage-level reservations, and network isolation all serve the same purpose: ensuring a demoted node cannot continue accepting writes because it has not yet learned it lost the election.

Applications need their own guard. Leases with monotonically increasing fencing tokens, validated at the storage layer, prevent a paused process from resuming and writing with stale authority.

Recovery is a decision, not a script

When divergence has already occurred, someone must decide which timeline survives. Prepare that decision in advance: identify the authoritative side by transaction volume, business criticality, or region of record, and keep tooling that can extract and replay the losing side's writes for review. Rehearse the procedure on a staging cluster, because the first attempt should never happen during an incident.

Key takeaways

  • Use odd voting counts and avoid symmetric two-site topologies
  • Place witnesses in an independent third failure domain
  • Pair quorum with real fencing and fencing tokens at the storage layer
  • Write and rehearse the divergence-reconciliation runbook in advance

Talk to KodeSync Resources

KodeSync Resources engineers distributed database synchronization, multi-master replication and high-availability data layers. Send us your environment and we will respond with a scoped audit plan.

Request system audit

Related articles