Log-based capture and why it wins
Query-based capture polls for updated timestamps and misses deletes, intermediate states, and anything written outside the expected pattern. Log-based capture reads the write-ahead log directly — logical decoding in Postgres, the binary log in MySQL, the oplog in MongoDB — and observes exactly what was committed, in commit order.
Debezium packages this as Kafka Connect source connectors, emitting before-and-after images with transaction metadata attached.
Topology decisions that matter early
A few choices are difficult to reverse once the pipeline carries production traffic.
- Key topics by primary key so compaction and ordering behave correctly
- Run a schema registry with compatibility rules enforced at write time
- Use the outbox pattern where event shape should not mirror table shape
- Size snapshot windows and consider incremental snapshots for large tables
Delivery guarantees and idempotency
CDC pipelines deliver at least once by default. Connector restarts replay from the last committed offset, so duplicates are normal and consumers must be idempotent — usually by upserting on primary key or tracking the log sequence number per record.
Ordering is guaranteed per partition, which means per key when keying is correct. Cross-table transactional ordering is not preserved unless the transaction metadata topic is consumed and buffered explicitly.
Operating the pipeline
The metric that matters most is replication slot lag on the source. An unconsumed slot retains WAL segments and will eventually exhaust disk on the primary — a CDC outage that becomes a database outage. Alert on slot lag and on connector task failure with equal urgency, and always maintain a documented path for dropping and rebuilding a slot with a re-snapshot.
Key takeaways
- Prefer log-based capture; polling misses deletes and intermediate states
- Key topics by primary key and enforce schema compatibility centrally
- Design consumers to be idempotent — delivery is at least once
- Alert on replication slot lag before it consumes primary disk
Talk to KodeSync Resources
KodeSync Resources engineers distributed database synchronization, multi-master replication and high-availability data layers. Send us your environment and we will respond with a scoped audit plan.
Request system audit