Overview
Apache Cassandra is a distributed wide-column NoSQL database built for write-heavy workloads, linear scaling and no single point of failure. This template deploys a Cassandra 5.0 cluster (cassandra:5.0) of replicas nodes in one location, each node on its own persistent volume, with password authentication enabled, a weekly nodetool repair cron job, and optional scheduled backups to AWS S3 or GCS.
Cluster credentials are not template values. Cassandra reads its superuser password, application role and keyspace from a dictionary secret you create before installing, so no password passes through Helm or lands in the release.
--gvc. Use a GVC with exactly one location: the template declares no per-location settings and forms a single-location cluster, so a GVC with several locations is not a supported shape for it.What Gets Created
Prerequisites
Onedictionary secret must exist before you install. It holds the credentials you type into every client and connection string. Secrets are org-level, so no GVC flag is involved.
Create the credentials secret
superuserPassword is the password of the built-in cassandra role, used for JMX and by the repair and backup jobs; username, password and keyspace are the application role and its keyspace, created on first boot. The keyspace value must be a valid unquoted CQL identifier (lowercase letters, digits and underscores):credentialsSecretName to this name. Secret names are org-wide, so give each release its own.Read the secret back later
-o yaml. A bare cpln secret reveal prints only a summary table, not the values:Installation
Create the prerequisite secret, then install with the name of that secret. The defaults give you three nodes of 1 CPU and 4 GiB each, a replication factor of 1, weekly repair and no backups:UI
CLI
Terraform
Pulumi
Configuration
Each block below is the shipped default for that part ofvalues.yaml.
Cluster Size and Replication
replicas— how many Cassandra nodes run. The token ring is split across them, so more nodes means more capacity and throughput.replicationFactor— how many copies of each partition the application keyspace keeps. With a factor of 3 every row exists on three nodes, and reads and writes atQUORUMconsistency succeed while one of them is down.
replicationFactor must not exceed replicas; the template refuses to render otherwise (replicationFactor (3) cannot exceed replicas (1)). The system_auth keyspace, which holds roles and passwords, is always replicated to every node regardless of replicationFactor.
replicationFactor: 3 with at least three replicas and read and write at QUORUM.Credentials
cpln://secret/... references, so the values appear in neither the Helm release nor the stored workload spec.
0 on the cluster’s first boot, and a flag file in its data directory (/var/lib/cassandra/.bootstrapped) stops them from being applied again. Changing the secret afterwards does not change the cluster — it only changes what clients and the cron jobs present, which then fails. To rotate on an existing cluster, follow Rotating Credentials.Image and Resources
cpu and memory apply to each node. Keep jvmHeapSize at roughly half of memory: Cassandra relies on off-heap memory for bloom filters, caches and the OS page cache. Choose clusterName before the first install — Cassandra stores it in the data directory and a node refuses to start when the configured name differs from the stored one.
Storage
general-purpose-ssd volume, mounted at /var/lib/cassandra and holding the data, commit log and saved caches. Capacities are in GiB; the volume grows by scalingFactor whenever free space drops below minFreePercentage, up to maxCapacity.
Multi-Zone
true, Control Plane spreads the nodes across availability zones within the location. Verify that your location supports multi-zone before enabling it.
SimpleStrategy, which ignores racks and zones when choosing where a partition’s replicas live, so a zone outage can take more than one replica of the same partition with it. Treat surviving a full zone outage as unverified with this template; the protection you can rely on is replicationFactor: 3 with QUORUM, which tolerates one node down.Internal Access
type accepts same-gvc, same-org or workload-list; for workload-list, list each allowed workload under workloads as an item of the form //gvc/GVC_NAME/workload/WORKLOAD_NAME.
workload-list and logical backups enabled, include the backup job itself — //gvc/GVC_NAME/workload/RELEASE_NAME-cassandra-backup — or its runs cannot reach the cluster. A change to internal_access can take a few minutes to take effect; re-test rather than trusting the first response.
Repair
nodetool repair reconciles the replicas. Repair must complete on every node at least once within gc_grace_seconds (10 days by default) or deleted data can reappear after a node recovers. The RELEASE_NAME-cassandra-repair cron job runs a full repair on each node in turn, authenticating over JMX with superuserPassword; the default weekly schedule fits inside the 10-day window with margin. Repair is resource-intensive on large datasets — schedule it for a low-traffic window rather than disabling it.
Backup
Connecting
UN (up, normal):
cqlsh from your own machine without opening the firewall, forward the port and connect to 127.0.0.1 with the application credentials:
Operations
Backing Up
Two modes are available;backup.type selects one. Neither mode saves the schema — keep your CREATE KEYSPACE and CREATE TABLE statements in version control, because both restore paths load data into tables that already exist. Backups keep no retention of their own; expire old folders with a bucket lifecycle rule.
AWS S3
- Create an S3 bucket. Set
backup.aws.bucketto its name andbackup.aws.regionto its region. - If you do not have a Control Plane Cloud Account for AWS, follow the Create a Cloud Account guide. Set
backup.aws.cloudAccountNameto its name. - Create an IAM policy with the following JSON, replacing
YOUR_BUCKET_NAME, and setbackup.aws.policyNameto its name. The template attaches this policy to the identity and nothing else, so the bucket in it is the only storage the backup can reach:
- Set
backup.aws.prefixto the folder path for the backups, then setbackup.enabled: true,backup.provider: awsand yourbackup.type.
GCS
- Create a GCS bucket. Set
backup.gcp.bucketto its name. - If you do not have a Control Plane Cloud Account for GCP, follow the Create a Cloud Account guide. Set
backup.gcp.cloudAccountNameto its name. - Add the Storage Admin role to the GCP service account associated with the Cloud Account. The template additionally binds the identity to
roles/storage.objectAdminon exactly the bucket inbackup.gcp.bucket. - Set
backup.gcp.prefixto the folder path for the backups, then setbackup.enabled: true,backup.provider: gcpand yourbackup.type.
Logical backup complete. Timestamp: ...; a physical run logs Physical backup complete. Tag: ... from each node:
Restoring a Backup
Both restore paths are/usr/local/bin/restore.sh in the backup image, driven by RESTORE_TIMESTAMP — the YYYY-MM-DDTHH-MM-SSZ folder name under your prefix. The keyspace and tables must already exist; the script loads data, it does not recreate schema.
restore.sh, but they have not been rehearsed against a live install of this template. Rehearse on a throwaway release before you depend on them.backup sidecar of every node in turn. It downloads that node’s SSTables onto the node’s data volume, loads them into the live tables with nodetool import — no restart — and deletes the download. Repeat for -1, -2 and so on; because each node owns a different token range, restoring only one node leaves the cluster with incomplete data:
PREFIX/TIMESTAMP/KEYSPACE/*.csv.gz and replays each file into its table with cqlsh COPY FROM; rows whose primary key is in the backup are overwritten, rows that are not are left in place. It needs the environment the RELEASE_NAME-cassandra-backup job runs with — BACKUP_PROVIDER, BACKUP_BUCKET, BACKUP_PREFIX, AWS_REGION for S3, and CASSANDRA_HOST, CASSANDRA_PORT, CASSANDRA_USER, CASSANDRA_PASSWORD, CASSANDRA_KEYSPACE — but that job is a cron workload with no long-running container to exec into. Running it means starting the backup image with that environment, bucket access and network access to the cluster, with RESTORE_TIMESTAMP set and /usr/local/bin/restore.sh as the command. This has not been exercised with the template.
Rotating Credentials
The secret is read only on the cluster’s first boot, so a rotation happens inside Cassandra first, and the secret and the running workloads are then brought in line. Do the three steps together: the nodes’ JMX password file is written from the secret at container start, and the repair and backup jobs authenticate with the secret’ssuperuserPassword, so a mismatch between the two makes those jobs fail until the redeploy.
Change the passwords inside Cassandra
cqlsh session as the superuser through the port-forward shown under Connecting (-u cassandra -p SUPERUSER-PASSWORD) and run whichever of these you are rotating:Update the secret
password and/or superuserPassword entries of your dictionary secret to the new values, so applications reading the secret and the cron jobs — which start a fresh container on every run — authenticate with them.Redeploy the nodes
Upgrading From 1.0.x
Template versions up to 1.0.1 took the cluster credentials as plain Helm values and shipped working defaults for them —supersecretpassword for the built-in cassandra role and password for the application role. Version 1.1.0 removed all four keys; the upgrade is in place and changes nothing on the data volumes.
Recover the credentials the cluster already uses
Create the secret with those same values
keyspaceName becomes the keyspace entry. Different values here do not change the cluster — they just leave clients and the cron jobs unable to authenticate.Remove the old keys and upgrade
credentialsSecretName, then upgrade.Rotate anything that came from a default
supersecretpassword and password were published in the public template repository, so treat them as compromised and follow Rotating Credentials.Upgrading From 1.1.0
Version 1.1.1 removesaws::ReadOnlyAccess from the backup identity; it only affects releases with backup.provider: aws. That AWS managed policy granted read access to every bucket in your AWS account and contained no write actions, so it never carried the backup itself — but it was silently supplying any read action your bucket-scoped policy happened to omit. Before upgrading, make sure your IAM policy contains the full action list under Backing Up; if it already matches, no action is needed. The identity now carries cpln-connector and your bucket-scoped policy only. Nothing else changes.
Scaling and Availability
Vertical. Raisecpu and memory together with jvmHeapSize (about half of memory); the data volumes grow on their own within volumes.data.autoscaling.
Adding nodes. Raise replicas and upgrade; the new nodes join the ring and the repair job and seed list follow the new count. New nodes do not stream existing data when they join — every node in this template is a seed, and Cassandra seed nodes do not auto-bootstrap. Once each new node shows UN in nodetool status, stream its share of the data onto it with nodetool rebuild, naming the cluster’s datacenter (the template sets it to the GVC location, for example aws-us-east-1), then run nodetool cleanup on every pre-existing node to drop the ranges they no longer own. This procedure follows the upstream topology changes guidance and has not been exercised against a live install of this template:
nodetool decommission) before it leaves the ring; lowering replicas without that is unverified and may lose data.
Node loss. Each node’s data lives on its own volume and comes back with it after a restart, and each node runs nodetool drain before it stops. While a node is down, the partitions for which it holds the only copy are unavailable — with the default replicationFactor: 1 that is every partition it owns. With replicationFactor: 3 and QUORUM, the cluster keeps serving reads and writes with one node down. Surviving a full zone outage is unverified; see Multi-Zone.
Troubleshooting
The workload never becomes ready and cpln logs is empty
The workload never becomes ready and cpln logs is empty
cpln helm install reported success, the deployment never reaches ready, and cpln logs returns no lines. cpln workload get-deployments RELEASE_NAME-cassandra --gvc GVC_NAME -o yaml shows under status.versions[].message:credentialsSecretName names a secret that does not exist (or was created in a different org).Fix: Create the secret as shown in Prerequisites, then run cpln workload force-redeployment RELEASE_NAME-cassandra --gvc GVC_NAME or wait several minutes for the deployment to recover on its own.First install sits at Waiting for all 3 replicas to join before bootstrapping
First install sits at Waiting for all 3 replicas to join before bootstrapping
0’s logs show Waiting for all 3 replicas to join before bootstrapping..., possibly followed by WARN: Timed out waiting for all replicas — proceeding with available nodes.Cause: Node 0 creates the roles and keyspace only after every node has joined the ring, so that token ranges are final first. It waits up to ten minutes before proceeding with whatever has joined.Fix: Wait. If the warning appears, check the other nodes’ logs with cpln logs '{gvc="GVC_NAME", workload="RELEASE_NAME-cassandra"}' --limit 50 --since 10m for why they did not join, and run nodetool status as shown under Connecting until every node reports UN.Install fails with replicationFactor cannot exceed replicas
Install fails with replicationFactor cannot exceed replicas
cpln helm install or upgrade exits with:replicas to at least replicationFactor and run the command again. Nothing was changed by the failed run.Upgrade fails with superuserPassword, username, password and keyspaceName were REMOVED
Upgrade fails with superuserPassword, username, password and keyspaceName were REMOVED
cpln helm upgrade exits with the render error quoted under Upgrading From 1.0.x and nothing is changed.Cause: Your values still set one of the four keys that 1.1.0 removed.Fix: Create the prerequisite secret with the cluster’s existing credentials, delete the four keys, set credentialsSecretName, and upgrade again.Clients are rejected with Bad credentials after the secret was changed
Clients are rejected with Bad credentials after the secret was changed
cqlsh or a driver fails with code=0100 [Bad credentials] message="Provided username USERNAME and/or password are incorrect" using the values currently in the secret.Cause: The secret is applied only on the cluster’s first boot. Editing it later changes what clients present, not the roles stored in Cassandra.Fix: Authenticate with the credentials the cluster was created with, then follow Rotating Credentials to bring Cassandra, the secret and the workloads back in line.The repair job logs Replica N repair failed
The repair job logs Replica N repair failed
RELEASE_NAME-cassandra-repair show WARN: Replica 1 repair failed. and end with ERROR: 1 replica(s) failed repair.Cause: The job connects to each node over JMX with the secret’s superuserPassword. The usual reason is that this value no longer matches what the nodes started with — a rotation that updated the secret without redeploying the nodes, or the reverse — or the node was down during the run.Fix: Confirm every node reports UN with nodetool status, then complete Rotating Credentials so the secret and the nodes agree, and wait for the next scheduled run.Backup run fails with AccessDenied on the S3 upload
Backup run fails with AccessDenied on the S3 upload
An error occurred (AccessDenied) when calling the PutObject operation.Cause: The IAM policy named in backup.aws.policyName is missing an action, or its Resource does not match backup.aws.bucket. From 1.1.1 nothing else on the identity fills the gap.Fix: Update the policy to the full action list under Backing Up with both the bucket ARN and BUCKET/*, then wait for the next scheduled run and check its logs.Important Notes
- Create the credentials secret before installing. A missing secret wedges the deployment with no log output; the only diagnostic is
status.versions[].messagefromcpln workload get-deployments. - Use a GVC with a single location. The template has no per-location settings and forms a single-location cluster.
- The default replication factor is 1. Set
replicationFactor: 3with at least threereplicasand useQUORUMbefore storing anything you cannot afford to lose. - Credentials apply only on first boot. Rotate inside Cassandra first, then update the secret and force a redeployment — in that order, together.
- Keep repair running. Every node must be repaired at least once every 10 days or deleted data can reappear.
- Backups do not include the schema. Keep your CQL schema in version control; both restore paths load data into existing tables.
- New nodes do not stream data by themselves. After raising
replicas, runnodetool rebuildon each new node andnodetool cleanupon the old ones. - Firewall changes are not instant. Allow a few minutes after changing
internal_accessbefore concluding a knob is broken.