Skip to main content

Overview

CDC Pipeline deploys a complete change data capture stack in a single release. It reimplements nothing: it bundles three catalog templates as dependencies — PostgreSQL Highly Available as the source database, Kafka as the event transport, and Debezium Server as the connector that reads the database’s change stream and publishes every row-level insert, update and delete to Kafka. This chart adds only the glue: one credentials secret for the database, the internal hostnames each component uses to reach the others, and checks at render time that the shared credentials agree.
This template deploys into an existing GVC that you already have. It does not create or manage a GVC. Every workload runs in each location the GVC has, so use a GVC with a single location.
Version 1.0.4 does not deliver change events with its default configuration. The bundled PostgreSQL cluster starts with wal_level = replica, while Debezium’s pgoutput decoding needs wal_level = logical; database.walLevel is checked at render but is not applied to the database. Kafka also enforces ACLs and the chart grants none to the debezium user. Both are described under Troubleshooting. Treat this release as a starting point you must finish configuring, and verify events arrive before you rely on it.

What Gets Created

Prerequisites

1

Create the Kafbat UI configuration secret

Kafbat UI is on by default and reads its whole configuration from an opaque secret mounted as a file, which must exist before you install. Replace ADMIN_PASSWORD with the value you will pass as kafka.kafka.listeners.client.sasl.admin.password and UI_PASSWORD with the login you want for the console. The full file format is in the Kafbat configuration reference.
Set kafka.kafbat_ui.configuration_secret to this name. If you do not want the console, install with --set kafka.kafbat_ui.enabled=false instead and skip this step.
If the secret is missing, the install still succeeds, cpln logs returns nothing for RELEASE_NAME-kafbat-ui, and the workload never becomes ready. The only place the missing secret is named is status.versions[].message:
Create the secret and the deployment recovers on its own after a few minutes, or force it with cpln workload force-redeployment RELEASE_NAME-kafbat-ui --gvc GVC_NAME.
2

Choose the passwords and the KRaft cluster ID

The database password, the Kafka passwords and the KRaft cluster ID are template values. Every default starts with change-me-cdc-pipeline-; those defaults are placeholders, not working secrets, so replace each one. Generate the KRaft cluster ID with the Kafka tooling, using the same image the brokers run:
The cluster ID is written into the Kafka log volumes at first start and cannot be changed afterwards without a new cluster.

Installation

Install into an existing GVC. The database password and the debezium Kafka password are each passed twice, because the database and Kafka read one copy and Debezium the other:

UI

Browse, install, and manage templates visually

CLI

Manage templates from your terminal

Terraform

Declare templates in your Terraform configurations

Pulumi

Declare templates in your Pulumi programs

Configuration

Each block shows the shipped defaults for one area of values.yaml. Only the keys you change need to go in your values file. The bundled components take their own values under postgres-highly-available, kafka and debezium-server; the linked component pages explain each option, but this release pins older component versions than those pages describe.

Database Credentials

The bundled database exists only to serve the pipeline, so its credentials are values and the chart writes them into the secret named by postgres-highly-available.config.credentialsSecretName. They must match the debezium-server.source.database values, which the chart checks at render.

PostgreSQL Highly Available

See PostgreSQL Highly Available for what each option does.

Kafka

The broker, listener, ACL and secret settings, plus the optional consoles. The exporter blocks (kafka.kafka_exporter, kafka.jmx_exporter) and kafka.kafka_client ship with working defaults; see Kafka for them.

Debezium Server

The connector’s source and sink. hostname and bootstrapServers are filled in from the release name when left empty, so the defaults already point at the bundled database and brokers.
Before the connector starts, its startup step creates the debezium_heartbeat table and, once wal_level is logical, the debezium replication slot in the source database, so the heartbeat needs no manual SQL. Pointing hostname or bootstrapServers at a database or Kafka outside this release still deploys the bundled components, and the render-time credential checks still apply. This release bundles Debezium Server chart 1.1.1, where credentials are values; the Debezium Server page documents a later version that reads them from a secret.

Connecting

Nothing in the pipeline is public except Kafbat UI. Everything else is reachable from inside the GVC. Read the Kafbat UI endpoint from the workload:
To exercise the pipeline end to end, first write to the source database. Tunnel to HAProxy from your own machine:
Then, in a second terminal, create a table and insert a row:
Next, read the change events from Kafka with the bundled client workload. Open a shell on it:
Inside it, write the admin credentials to a client properties file, list the topics and consume the table’s topic:
If no dbserver1. topic appears, read the connector log and see Troubleshooting:
Applications in the GVC consume the same way, with these client properties and the bootstrap address above. To give a consumer its own Kafka user rather than admin, append it to the comma-separated kafka.kafka.listeners.client.sasl.users list, with its password at the same position in passwords, and give it access, either with Kafka ACLs or by listing it in kafka.kafka.acl.superUsers.

Operations

Backing Up

Every volume set in the pipeline takes a final snapshot when it is deleted and keeps it for seven days; the Kafka log volume sets also take a snapshot daily at 00:00 UTC. Take extra snapshots before a risky change. The source database is the system of record:
The Debezium offsets on RELEASE_NAME-debezium-data record how far the connector has read; losing them makes it snapshot the source again from the start:
The bundled PostgreSQL template’s own backups (pg_dumpall or WAL-G to object storage) are switched off here with postgres-highly-available.backup.enabled: false. Turning them on through this chart has not been verified; see PostgreSQL Highly Available for how they work in the standalone template.

Restoring a Backup

Restoring any of this pipeline’s volume snapshots has not been verified. A restored database whose replication slot no longer matches the Debezium offsets, or Kafka brokers whose metadata no longer matches their peers, may not resume cleanly. Rely on PostgreSQL and Kafka replication for the loss of a single member, and treat snapshots as a last resort.

Upgrading from 1.0.1

Versions 1.0.0 and 1.0.1 bundle a PostgreSQL Highly Available release that does not compact its etcd cluster, so etcd’s backend grows with time alone until it reaches its 2 GiB quota and goes read-only, taking PostgreSQL failover with it. Every install of those versions is affected, because the database cannot be switched off. Upgrading to 1.0.2 or later turns compaction on and stops further growth, but it cannot shrink a backend that has already grown, and a cluster that has already raised a NOSPACE alarm needs operator recovery rather than an upgrade. The mechanism and the checks are on the PostgreSQL Highly Available page under etcd History Compaction.

Upgrading from 1.0.3

Version 1.0.4 moved the database credentials. They used to be postgres-highly-available.postgres.username, .password, .database and .walLevel; they are now database.username, database.password, database.name and database.walLevel, and this chart creates the credentials secret named by postgres-highly-available.config.credentialsSecretName. Old keys left under postgres-highly-available.postgres are no longer read. An existing database keeps the credentials it was created with, so set database.* to those exact values; different values leave the members unable to authenticate after their next restart. Recover them before upgrading, from your values file or from the secret the earlier version created:
This upgrade path has not been verified end to end. Snapshot the database volume set first, as shown under Backing Up.

Scaling and Availability

  • The source database runs three Patroni members behind HAProxy and fails over automatically when the leader is lost; Debezium connects through HAProxy, which routes to the current leader, and its startup step creates the replication slot.
  • Kafka runs three brokers with a replication factor of 3. kafka.kafka.replicas accepts 1, 3, 4 or 5; see Kafka for the rules.
  • Debezium Server is a single replica and cannot be scaled out.
  • Keep postgres-highly-available.etcd.replicas odd, and choose it at install time.

Troubleshooting

Cause: in version 1.0.4 this is expected with the defaults, for up to three reasons. The bundled PostgreSQL cluster runs with wal_level = replica, which pgoutput logical decoding cannot use; database.walLevel is validated but not applied. Kafka enforces ACLs with allowEveryoneIfNoAclFound: false, and the debezium user is neither a super user nor granted any ACL. Or debezium-server.sink.kafka.saslPassword differs from the debezium entry in kafka.kafka.listeners.client.sasl.passwords — the render-time check compares only the user name.Fix: read the connector log, filtered server-side, to see which applies:
Align the two Kafka passwords if they differ. Adding User:debezium to kafka.kafka.acl.superUsers puts it in the brokers’ super-user list. Patroni keeps wal_level in its cluster-wide configuration, so it is not changed by any value of this chart. Neither change has been verified end to end on this template.
Cause: the secret named in kafka.kafbat_ui.configuration_secret did not exist when the workload was created. The container never starts, so there is no log output.Fix: read status.versions[].message from cpln workload get-deployments RELEASE_NAME-kafbat-ui --gvc GVC_NAME -o yaml to see which secret is missing, create it as shown in Prerequisites, then wait a few minutes or run cpln workload force-redeployment RELEASE_NAME-kafbat-ui --gvc GVC_NAME.
Cause: database.username, database.password or database.name differs from debezium-server.source.database.user, .password or .name. The error text still calls the database side postgres-highly-available.postgres, the name these keys had before 1.0.4.Fix: set both sides to the same values; the install line in Installation passes the password twice for this reason.
Cause: debezium-server.sink.kafka.saslUsername does not appear in kafka.kafka.listeners.client.sasl.users.Fix: add the user to the Kafka list, or set saslUsername to a user already in it.
Cause: database.walLevel was set to something other than logical.Fix: leave it at logical.
Cause: the database credentials secret is named my-cdc-pipeline-db-credentials by default, without the release name, and secrets are org-wide.Fix: give every release in the org its own postgres-highly-available.config.credentialsSecretName.
Cause: a replication slot is not advancing — the connector stopped reading, or a slot was left behind after debezium-server.source.postgres.slotName changed. PostgreSQL keeps WAL for an inactive slot indefinitely.Fix: drop the abandoned slot on the source, using its name:

Important Notes

  • In version 1.0.4 the pipeline does not deliver change events with its defaults; see Troubleshooting and verify events arrive before relying on it.
  • Replace every change-me-cdc-pipeline- value before installing; they are plain template values and land in the Helm release.
  • Pass the database password and the debezium Kafka password to both the component that owns them and debezium-server; only the database values and the Kafka user name are checked at render.
  • kafka.kafka.secrets.kraft_cluster_id is fixed once the brokers have started.
  • Kafbat UI is open to the internet by default; set a strong console login or narrow kafka.kafbat_ui.firewall.external_inboundAllowCIDR.
  • Give each release its own postgres-highly-available.config.credentialsSecretName; the default name is shared org-wide.
  • Uninstalling deletes every volume set — the source database, the Kafka topics and the Debezium offsets; final snapshots are kept for seven days.
  • Component versions are pinned in this chart (PostgreSQL Highly Available 2.5.0, Kafka 4.0.1, Debezium Server 1.1.1) and change only with a new CDC Pipeline version.

External References