← Back to Blog
Managed CDC vs self-hosted Debezium
CRM & ERP Integration

Managed CDC vs self-hosted Debezium

byBruno Galo · Published on 05 Oct 2026

Available inCatalàEnglishEspañolPortuguês

Short answer. Managed CDC usually wins when nobody on your team runs Kafka in production today, when you have only a handful of pipelines, or when you have a deadline. Self-hosting wins when you already operate Kafka well, when network or residency constraints rule out a vendor, or when volume is high and steady enough that usage-based pricing grows faster than a platform engineer's time. In between sits a middle path: managed Kafka with your own Debezium connectors, or Debezium Server without Kafka. If what you actually need is two-way operational sync between a CRM and a database, neither side of this decision is the whole answer.

The licence is the smallest line. Self-hosted Debezium is free software running on infrastructure and people you pay for. Managed CDC is a bill that replaces part of that people-time. This article gives you a cost model you can fill in with your own numbers, the failure modes that usually decide the question, and a short answer for each team shape. We write it as engineers who run both: we are a Stacksync partner for managed operational sync, and we build custom CDC pipelines when that is the better fit.

What "managed CDC" and "self-hosted CDC" mean in 2026

Change data capture (CDC) reads row-level changes from a database log and streams them somewhere else. The choice is not two boxes, managed or not. It is a spectrum, and where you sit on it decides what you operate.

Option What you run What the vendor runs Examples
Fully self-hosted Debezium connectors, Kafka Connect, Kafka Nothing Debezium on your own Kafka 4.x cluster (KRaft mode, no ZooKeeper)
Debezium without Kafka Debezium Server or the embedded Debezium Engine Nothing Streaming to sinks such as Kinesis, Pub/Sub, Redis or HTTP
Managed Kafka, your connectors Debezium connector choice and configuration Connect workers, brokers Amazon MSK Connect, Aiven for Apache Kafka Connect
Fully managed CDC connector Source database settings, downstream consumers Connector runtime Confluent Cloud PostgreSQL CDC Source V2 (Debezium)
CDC inside a sync or ELT product Source and target configuration The pipeline end to end Fivetran, Airbyte, Estuary-type products, Stacksync for two-way operational sync

These are examples, not a ranking. Check every product claim against the vendor's own documentation before you decide, because plans and connectors change.

Two facts keep the self-hosted end of the spectrum less heavy than it used to be. Kafka 4.x runs without ZooKeeper. And Debezium 3.7, released on 29 September 2026, shipped the first official Debezium CLI and let Debezium Platform deploy to hosts over SSH instead of only Kubernetes. That is less operational work than before. It is still yours.

At the managed end, Confluent Cloud's PostgreSQL CDC Source V2 is a Debezium-based connector billed per connector task-hour plus data transfer. The pricing page lists some premium connectors as "contact us", so we do not quote a rate here.

The cost model: every line, not just the licence

Most spreadsheets compare the licence and the cloud bill. The lines that decide the answer are the ones nobody enters.

Cost line Self-hosted Managed
Software licence 0 (Apache 2.0) Subscription or usage
Compute for Connect workers and Kafka brokers Your cloud bill In the price, or partly (MSK-style)
Storage and retention (topics, WAL held by slots) You size it Usage-based
Network, egress, cross-AZ traffic Yours Often billed separately
Schema registry Run it or buy it Usually included or an add-on
Monitoring and alerting (lag, slot size, task failures) You build it Partly included
Upgrades (Debezium, Kafka, JVM, database versions) Your sprint time Vendor
On-call Your rota Vendor for the platform, you for data issues
Incident recovery (re-snapshot, replay) Your engineers Shared
Knowledge concentration (bus factor) High risk Lower
Exit cost Low (open source) Migration effort

Then put the lines into one formula:

Annual TCO = infrastructure + licences + (engineer hours per month × loaded hourly cost × 12) + incident allowance

Plug in your own loaded cost and your own cloud bill. We deliberately give no euro figures: any price we could print would be either invented or out of date, and the interesting number is yours. The same goes for the crossover. Self-hosting gets cheaper as volume grows only if the people-time stays flat, so compute the point where usage pricing overtakes your engineers' time with your own inputs, and treat any fixed threshold you read online, including in vendor blogs, with suspicion. If you want the Salesforce-specific version of this argument, read the Heroku Connect build-vs-buy breakdown.

We will fill in this model with you: book a 30-minute call and bring your cloud bill and your on-call rota.

What breaks in production (and who fixes it)

The honest comparison is in the failure modes. For each one, the question is what the vendor handles and what still lands on you.

PostgreSQL replication slots holding WAL

If the connector stops consuming, the replication slot keeps WAL segments and the disk on your primary fills. Mitigations are setting max_slot_wal_keep_size (PostgreSQL 13 and later), alerting on slot lag and using a heartbeat on quiet databases. Self-hosted: you do all of it. Managed: the vendor runs the connector, but the slot lives on your database, so the risk and the alerting remain yours. Airbyte's own documentation notes that a slot invalidated after max_slot_wal_keep_size is exceeded needs recreating and a full re-sync.

Schema changes

DDL on the source table changes what the connector emits. Schema registry compatibility rules decide whether downstream consumers break. Self-hosted: you design and enforce the compatibility policy. Managed: the platform may include a registry, but your teams still own the contracts with each consumer.

Snapshots and re-snapshots

An initial snapshot reads the source tables and loads the source database. Incremental snapshots reduce that cost, and a long outage can force you back into one. During a long snapshot the slot is not advanced, so on a large table WAL accumulates. Self-hosted: you schedule and watch it. Managed: the vendor runs the snapshot, you still decide when the load on your database is acceptable.

Delivery semantics

Plan for at-least-once delivery. Consumers must be idempotent, and we would not promise end-to-end exactly-once for any option. Self-hosted: you design the idempotency. Managed: the same, because it lives in your consumers.

Connector task failures and restarts

Tasks fail, restart and sometimes hit a poison message. Dead-letter queues and clear runbooks matter more than the platform. Self-hosted: you carry the pager. Managed: the vendor restarts workers, you still triage bad data.

Upgrades

Debezium ships frequent 3.x releases, Kafka has major versions, and a major upgrade of your managed PostgreSQL needs planning around logical replication slots. Self-hosted: upgrade work is sprint work. Managed: the vendor upgrades the platform, you plan the database side.

When self-hosting is the right call

Self-hosting is the right call in these situations:

  • You already have a Kafka platform team, so the marginal cost of one more connector is small.
  • A strict VPC, on-premises or residency requirement that no vendor can meet.
  • High, steady volume where usage pricing outgrows people cost. Compute the crossover with your own numbers.
  • You need custom single message transforms or connectors the managed catalogue does not offer.
  • Owning the pipeline is strategic, for example because it is part of your product.

When managed wins

Managed wins in these situations:

  • You have no Kafka in production today.
  • One person is the only one who understands the Connect cluster. There is no magic team size: the test is whether that person can go on holiday.
  • The migration is deadline-driven, for example leaving Heroku Connect.
  • You have many small pipelines, each too small to justify its own care.
  • Audit or compliance work prefers vendor SOC 2 or ISO evidence.

The European angle: data residency and GDPR

A CDC stream copies personal data, so the pipeline is in scope for your GDPR records of processing. Check the vendor's regions, whether an EU region is available for the connector you need, the list of sub-processors and the data processing agreement. Most major managed vendors offer EU regions, but verify per vendor and per product. Self-hosting inside your own EU cloud region is the simplest residency story. This is not legal advice.

Migrating from self-hosted to managed Kafka (or back)

Teams searching for the best managed Kafka service for a migration are usually asking how to move without a re-snapshot, not which vendor to pick. We do not rank vendors. The sequence matters more than the logo:

  1. Inventory topics, connectors, configurations and committed offsets.
  2. Choose the mirroring route, for example MirrorMaker 2 or a vendor's cluster-linking-style option, and test it on a non-critical topic.
  3. Keep replication slot continuity so the connector resumes where it stopped and does not fall back into a snapshot of your production tables.
  4. Plan cutover and rollback: who switches consumers, in what order, and what is the point beyond which you cannot go back.
  5. Run both in parallel long enough to compare lag and counts before you decommission anything.

The reverse move, from managed back to self-hosted, follows the same steps. Exit cost is low for open-source Debezium and higher for a managed platform, which is one more cost line to record.

What about two-way sync (CRM and database)?

CDC is a one-way stream of changes. Writing back, for example between Salesforce and PostgreSQL, needs conflict handling, idempotency and loop prevention. Those layers are the hard part and are explained in the build-vs-buy breakdown. If that is your case, go to our Salesforce–PostgreSQL integration page and the Stacksync partner page.

Decision checklist

Answer yes or no. The pattern, not the count, is what matters.

  1. Do you run Kafka in production today? (Yes points to self-hosted or middle path, no points to managed.)
  2. Is there more than one person who can fix the Connect cluster at 3 a.m.?
  3. Will you have more than a handful of pipelines in the next 12 months?
  4. Is there a residency or network constraint that a vendor cannot meet?
  5. Are your source databases all ones the managed option supports?
  6. Do you need custom transforms or connectors?
  7. Do you need two-way sync? (Yes: neither option alone, see the previous section.)
  8. Is the acceptable lag measured in seconds, or is minutes fine?
  9. Does the budget owner prefer an operating subscription or engineer time?
  10. Is there a hard deadline within the next quarter?

Mostly "no" on the first three and "yes" on the last: managed. Mostly "yes" on the first three and on 4 or 6: self-hosted. A mix: the middle path of managed Kafka with your own connectors, or Debezium Server without Kafka.

Frequently asked questions

Is it worth paying for managed CDC instead of running your own Kafka Connect?

Usually yes if nobody on your team already runs Kafka in production, or if one person holds all the knowledge. Self-hosting pays off when you already operate Kafka well, have residency constraints, or have high steady volume where usage pricing outgrows engineer time.

What does self-hosted Debezium really cost?

The software is free under Apache 2.0. The cost is infrastructure (Kafka brokers, Connect workers, storage, network) plus engineering time for monitoring, upgrades, on-call and incident recovery. Model every line with your own numbers.

Which CDC platform has the lowest total cost of ownership?

It depends on volume, team and existing infrastructure. No platform is cheapest for everyone. Compare the full cost lines, not the licence.

Do I need Kafka to use Debezium?

No. Debezium Server and the embedded Debezium Engine stream changes to other targets without Kafka. Kafka Connect is still the most common deployment.

Does managed CDC remove the replication-slot risk on PostgreSQL?

No. The slot lives on your database. If the connector stops consuming, WAL accumulates. Set max_slot_wal_keep_size and alert on slot lag either way.

Debezium vs Airbyte for CDC?

Debezium is a CDC engine that streams row changes. Airbyte is an ELT platform that can use CDC for some sources, and its Postgres CDC relies on Debezium under the hood. Choose by the latency you need and by what you want to operate. The build-vs-buy post compares them for the Salesforce case.

Can managed CDC run in an EU region for GDPR?

Most major vendors offer EU regions. Check region, sub-processors and the data processing agreement per vendor. Self-hosting in your own EU cloud region is the simplest residency story.

Is CDC enough for two-way sync between Salesforce and PostgreSQL?

No. CDC is one-way. Two-way sync needs write-back, conflict rules and loop prevention. See our Salesforce–PostgreSQL page.

Closing: not sure which side of the line you are on?

Start by listing your pipelines, who is on call for each and what a day of lag costs. That list usually settles the question faster than a vendor comparison. If you want a second opinion from people who run both, book a 30-minute call. For the wider cluster, read the Heroku Connect end-of-life migration guide, the CDC primer on Salesforce and Heroku Connect and iPaaS vs point-to-point vs middleware.

About the author

Bruno Galo is the founder of Atypical Tech, a NetSuite consultancy serving mid-market clients across Iberia. He specializes in connecting CRM and ERP systems for seamless order-to-cash workflows, building automated order management pipelines that eliminate manual data entry between sales and finance teams. As an official Stacksync implementation partner, Bruno designs and deploys AI agents on integration platforms to handle exception routing, document processing, and reconciliation — turning fragmented order flows into reliable, self-monitoring systems.

LinkedIn: https://www.linkedin.com/in/brunogd

Sources

Comments

No comments yet.

Leave a comment

Your comment will be reviewed before publishing.

An unhandled error has occurred. Reload 🗙