Overview
DuckDB is an in-process analytical SQL engine that reads and writes Parquet, CSV and JSON directly, queries object storage over the S3 API, and can attach live PostgreSQL, MySQL and SQLite databases. This template runs a SQL script of yours on a cron schedule and exits: it is a batch job runner, billed only for the minutes a run is executing, with nothing running between runs.What Gets Created
/tmp.
Prerequisites
None for a default install. The shippedsql.inline is a self-test that needs no credentials, no bucket and no cloud account. Every run does download its extensions from extensions.duckdb.org:443; the template allows all outbound traffic, so this works unless your organization restricts egress.
A real job has to write its results somewhere durable — object storage or an attached database — and the features that make that possible each need something created before install:
Object storage (when objectStore.type is aws or s3-compatible)
none registers a DuckDB S3 secret so your SQL can read and write s3:// paths.- AWS S3 (keyless)
- S3-compatible
- Create your S3 bucket and note its region.
- If you do not have one, create a cloud account for your AWS account.
- Create an IAM policy scoped to that bucket (replace
YOUR_BUCKET_NAME). Drops3:PutObjectands3:DeleteObjectif the job only reads.
objectStore.type: aws, objectStore.aws.region to the bucket’s region, objectStore.aws.cloudAccountName to the cloud account’s name, and objectStore.aws.policyName to the IAM policy’s name. The chart refuses to render if any of the three is blank.A script secret (when using sql.secretName instead of sql.inline)
--encoding plain; its payload is mounted as the job script:sql.secretName to this name. The template’s own script secret is then not created at all.Installation
A default install needs no--set values — it runs the built-in self-test script on the default schedule:
UI
CLI
Terraform
Pulumi
Configuration
Image and Schedule
SQL Script
sql.secretName wins when both are set, and is the right choice when the script is too large for a values file or you want to change it without a Helm upgrade. The chart refuses to render when both are empty.
Before your script, the job runs a template-generated preamble containing .bail on and the derived memory_limit, threads, temp_directory and extension_directory settings, then the object-store secret when one is configured. Because the preamble runs first, a SET in your own script always wins. .bail on stops execution at the first SQL error.
Credentials in SQL
getenv('NAME'). Only a cpln://secret/... reference appears in the rendered workload; the value itself never lands in your values file. Names must be UPPER_SNAKE_CASE and unique, and AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN and AWS_DEFAULT_REGION are rejected because objectStore owns them.
To attach a database, name the entry after the driver’s own password variable and leave the password out of the connection string — the driver reads it from the environment. ATTACH takes a string literal, so getenv() cannot be concatenated into it (see Troubleshooting).
RELEASE_NAME-postgres is the workload of a postgres template installed in the same GVC. getenv('NAME') still works anywhere an ordinary expression is allowed — a WHERE clause, a computed column, a COPY destination built with ||.
Resources and Tuning
minMemory may not exceed maxMemory, minCpu may not exceed maxCpu, maxCpu:minCpu may not exceed 4:1, and memoryLimitPercent must be between 20 and 80.
Every job must fit in memory. No volume is attached, so size the work with resources.maxMemory and tuning.memoryLimitPercent rather than relying on spill. DuckDB’s out-of-core operators still function, but they spill to container-local scratch at /tmp/duckdb-temp, bounded by container disk, not to a sized, persistent volume. Extensions are written to /tmp/duckdb-extensions, which is likewise container-local, so they are downloaded again on every run; httpfs and aws autoload on first use of an s3:// path, so most scripts never issue an explicit INSTALL.
Object Storage
objectStore.type decides whether the preamble registers a DuckDB S3 secret; anything other than none lets your SQL read and write s3:// paths directly. The bucket, cloud account, IAM policy and credentials secret each mode needs are set up under Prerequisites. With s3-compatible, the identity is granted reveal on exactly the credentials secret you name, and the two keys are injected as AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY.
Connecting
This template exposes nothing to connect to — it is a job, not a server. Observe and drive it instead:Operations
Confirming a Run Succeeded
A run succeeded only when both of these are true:- The line
duckdb-job-completeappears in the job’s log output. - The run’s status is
Successfulincpln workload cron get RELEASE_NAME-duckdb --gvc GVC_NAME.
- The marker alone is not enough. A run terminated for exceeding
activeDeadlineSecondscan finish its script inside the termination lag and print the marker on its way out, while the platform has already recorded the run as failed. - The status alone is not enough. DuckDB’s CLI can exit
0on some failed scripts (upstream issue #16574), which would leave the run looking successful.
.bail on stops execution at the first error — so a SQL error aborts the script and the marker is never printed.
Scaling and Availability
DuckDB is one process over one database, with no clustering and no failover in any edition, so there is deliberately noreplicas knob: the workload is pinned to one replica per GVC location. Installing the template several times — different scripts, schedules or GVCs — is real scale-out, but it is not high availability. If a run’s container dies, that run did not happen, and a failed run is not retried; the next scheduled run starts normally.
Nothing is stateful, so uninstalling and reinstalling is free: cpln helm uninstall RELEASE_NAME --gvc GVC_NAME removes everything the template created and leaves your prerequisite secrets in place.
Troubleshooting
Nothing is listening after install
Nothing is listening after install
cpln workload cron get RELEASE_NAME-duckdb --gvc GVC_NAME and output with cpln logs. If you need an always-on SQL endpoint, install trino.Parser Error: syntax error at or near "||"
Parser Error: syntax error at or near "||"
ATTACH statement that builds its connection string with || getenv('PGPASSWORD') fails:ATTACH takes a string literal, not an expression. This applies equally to the postgres, mysql and sqlite attach types.Fix: name the secretEnv entry after the driver’s own password variable (PGPASSWORD for postgres, MYSQL_PWD for mysql) and leave password= out of the connection string. If a credential truly must sit inline, put the whole connection string into a sql.secretName secret instead.The run was killed with no SQL error, or the container reports OOMKilled
The run was killed with no SQL error, or the container reports OOMKilled
cpln workload get-deployments RELEASE_NAME-duckdb --gvc GVC_NAME -o yaml shows an OOMKilled restart reason.Cause: the job needed more memory than the container has. The most common trigger is a SET memory_limit in your own script that exceeds what resources.maxMemory can back — your SET overrides the template’s derived limit.Fix: remove the SET memory_limit from your script and change tuning.memoryLimitPercent instead, or raise resources.maxMemory. Do not plan around spill; every job must fit in memory.A run that worked yesterday now fails installing an extension
A run that worked yesterday now fails installing an extension
s3:// read or at an explicit INSTALL, with an HTTP or download error naming extensions.duckdb.org.Cause: there is no cache volume, so every run downloads its extensions from extensions.duckdb.org:443. The template allows all outbound traffic, so this usually means egress to that host was blocked at the organization or network level after the job was set up.Fix: restore outbound access to extensions.duckdb.org:443. The dependency is per run and cannot be removed by this template.The run status is failed but the log shows duckdb-job-complete
The run status is failed but the log shows duckdb-job-complete
cpln workload cron get records the run as failed while the completion marker is in the output.Cause: the run exceeded activeDeadlineSeconds. The platform records the failure when the deadline passes, but terminates the container up to roughly two minutes later, and the script finished inside that lag.Fix: raise activeDeadlineSeconds if the run time is legitimate, or reduce the work. Treat this run as failed when deciding whether its output is complete.helm install is refused with a message starting duckdb:
helm install is refused with a message starting duckdb:
duckdb: ... message naming a value.Cause: the chart validates its values before anything is created: the schedule must have five fields; tuning.memoryLimitPercent must be 20–80; minMemory/minCpu may not exceed their max counterparts; maxCpu:minCpu may not exceed 4:1; an objectStore.s3Compatible.endpoint may not carry a scheme; secretEnv[].name must be UPPER_SNAKE_CASE, unique and not one of the four AWS_* names; and the required fields for the chosen objectStore.type must be set.Fix: change the value the message names and install again. Nothing was created, so there is nothing to clean up.The job never runs and cpln logs returns nothing
The job never runs and cpln logs returns nothing
cpln logs shows no lines at all, and a manual cpln workload cron start produces no output either.Cause: a secret named in sql.secretName, secretEnv[] or objectStore.s3Compatible.credentialsSecretName does not exist, so the container cannot start.Fix: read status.versions[].message in cpln workload get-deployments RELEASE_NAME-duckdb --gvc GVC_NAME -o yaml, create the missing secret, then run cpln workload force-redeployment RELEASE_NAME-duckdb --gvc GVC_NAME.Important Notes
- This is a batch job, not a query service. Nothing is listening after install; that is correct behavior. For an always-on SQL endpoint, use trino.
- Check both success signals — the
duckdb-job-completemarker and a run status ofSuccessful. Neither one alone is reliable. activeDeadlineSecondsterminates a run within roughly two minutes of the deadline, not exactly at it, and the container bills for that window.- Every job must fit in memory. No volume is attached, so raise
resources.maxMemoryrather than relying on spill. - Never
SET memory_limithigher than the container. Changetuning.memoryLimitPercentinstead. - No
.duckdbfile is kept. Every run starts from an empty in-memory database; write results to object storage or an attached database. - Extensions are downloaded on every run from
extensions.duckdb.org:443. Every execution depends on that host being reachable. - Prerequisite secrets must exist before install, and uninstalling the release does not delete them — the template only removes the secrets it created itself.
- A cron workload runs in every location of its GVC. Install into a single-location GVC unless you want one run per location per schedule.
- Installing this template several times is scale-out, not high availability. There is no failover and a failed run is not retried.
- One script per install, by design. Multi-step, conditional or retrying pipelines belong in airflow.