Skip to main content

Overview

Spark deploys an Apache Spark standalone cluster — a distributed engine for batch ETL, large-scale data transformation, and SQL/DataFrame analytics. The template provisions a master (cluster manager and Web UI), a horizontally-scalable worker tier, an optional Spark Connect gRPC server for thin/remote clients, and an optional History Server that reads completed applications’ event logs back from an S3 or GCS bucket.

Architecture

  • Master — cluster manager and Web UI on :8080, cluster RPC on :7077. Fixed single replica.
  • Worker — the executor host tier, with its own Web UI on :8081. workers.replicas (default 1); more than one forms a multi-worker cluster.
  • Spark Connect (optional) — a gRPC server on :15002 for thin/remote client submission. Off by default.
  • History Server (optional) — a Web UI on :18080 that lists completed applications by reading their event logs back from object storage. Off by default; requires an S3/GCS bucket.

What Gets Created

  • Standard Spark Master Workload — {release}-spark-master, the cluster manager and Web UI. One replica; UI port :8080 (http), RPC port :7077 (tcp).
  • Standard Spark Worker Workload — {release}-spark-worker, the executor host tier. workers.replicas replicas; Web UI :8081 and RPC :7078 (tcp, internal only).
  • Standard Spark Connect Workload (optional) — {release}-spark-connect, a gRPC server on :15002 for remote client submission. Created only when connect.enabled: true.
  • Standard History Server Workload (optional) — {release}-spark-history, the completed-application dashboard on :18080. Created only when historyServer.enabled: true.
  • Config Secret — {release}-spark-conf, an opaque secret (encoding: plain) holding the rendered spark-defaults.conf (reverse-proxy, pinned RPC ports, and event-log/S3A settings). Always created.
  • Identity & Policy — one shared identity with a policy granting reveal on the config secret. When the History Server is enabled, the identity also carries the cloud-account binding for keyless object-storage access.
This template does not create a GVC. You must deploy it into an existing GVC. No persistent volume is created either — worker scratch is ephemeral, and the History Server’s durable store is the object-storage bucket.

Prerequisites

  • Default install: none. The master and a single worker come up with no external dependencies.
  • History Server (historyServer.enabled: true): an object-storage bucket and a Control Plane cloud account — AWS S3 or GCS. See Storage Setup.
Install the template using your preferred method:

UI

Browse, install, and manage templates visually

CLI

Manage templates from your terminal

Terraform

Declare templates in your Terraform configurations

Pulumi

Declare templates in your Pulumi programs

Configuration

The default values.yaml for this template:

Image

  • image — The Spark container image used by every tier. The chart is shipped and tested on apache/spark:4.0.4-scala2.13-java17-python3-ubuntu.

Resources

Each tier (master, workers, connect, historyServer) exposes its own minCpu / maxCpu / minMemory / maxMemory. maxCpu maps to the workload’s CPU limit and maxMemory to its memory limit, with minCpu / minMemory as the reservation.

Workers

  • workers.replicas — Number of worker replicas. 1 is the proven single shape; a value greater than 1 forms a multi-worker cluster and tasks are spread across replicas.
  • workers.cores — Cores each worker offers to the cluster (SPARK_WORKER_CORES).
  • workers.memory — Memory each worker offers to executors (SPARK_WORKER_MEMORY). This is Spark’s offer to executors, not the container limit — keep it below workers.maxMemory.

Spark Connect

  • connect.enabled — When true, starts a Spark Connect gRPC server on :15002 for thin/remote client submission. Off by default.
A PySpark Connect client needs its own dependencies. The apache/spark image ships only the server side — install pyspark[connect] (or pandas, pyarrow, and grpcio) in the client environment.

History Server

  • historyServer.enabled — When true, starts the History Server on :18080, which lists completed applications by reading their event logs back from object storage. Off by default; it requires storage.* and a real bucket. See Storage Setup.

Object Storage

Configured only when the History Server is enabled. Access is keyless (UCI) — the workload identity federates with your cloud account and no access keys are stored in values. The cluster’s driver tiers write event logs to this bucket and the History Server reads them back. See Storage Setup for the per-provider steps.

Extra Spark Configuration

  • sparkDefaults — A map of key: value entries merged verbatim into spark-defaults.conf, for any Spark setting the template does not expose directly (for example spark.sql.shuffle.partitions). See the Spark configuration reference.

Access

  • internalAccess.type — Controls which workloads may reach the cluster over the GVC network:
internalAccess.type: workload-list must include the cluster’s own workloads — the master, workers, and Connect server address each other over the GVC network, so an incomplete list breaks the cluster itself. same-gvc (the default) avoids this.
  • publicAccess.enabled — When true, exposes the Master and History Server Web UIs on a public *.cpln.app canonical endpoint. The worker and Connect internal UIs stay internal and are reached through the master’s reverse proxy. See Important Notes — the UIs are unauthenticated.

Submitting Jobs

Submit from the master container or any client workload in the GVC. A driver must advertise its pod IP — set spark.driver.host, or executors cannot connect back to the driver and the job stalls relaunching executors:
A successful SparkPi run prints Pi is roughly 3.14.... For remote/thin-client submission without an exec shell, enable Spark Connect and point a client at sc://…:15002.

Connecting

All endpoints are internal GVC addresses. Replace RELEASE_NAME and GVC_NAME with your values. The worker and running-application (:4040) UIs are not exposed directly — the master serves them through its reverse proxy at /proxy/{id}/, so they are reachable through the master endpoint.

Reaching a Private UI

With publicAccess: false (the default), reach any Web UI in a browser with cpln port-forward:
Then open http://localhost:8080. The History Server is reached the same way — cpln port-forward RELEASE_NAME-spark-history 18080:18080 --gvc GVC_NAME. There are no credentials: Spark’s Web UIs and cluster port are unauthenticated.

Storage Setup

The History Server reads completed applications’ event logs from a bucket that the cluster’s driver tiers also write to. Access is keyless — the workload identity federates with your cloud account, so no keys are stored in values. Set historyServer.enabled: true and configure storage.*.
1

Create a bucket

Create your S3 bucket and set storage.bucket and storage.region to match. Set storage.provider: aws.
2

Set up a Cloud Account

If you do not have one, create a Cloud Account for your AWS account. Set storage.cloudAccountName to its name.
3

Create a bucket-scoped IAM policy

Create an AWS IAM policy with the JSON below (replace YOUR_BUCKET_NAME), then set storage.policyName to the policy’s name:
4

Set the prefix

Set storage.prefix to the folder within the bucket where event logs are written.
The History Server crash-loops if the bucket does not exist. Enable it only with a real bucket and cloud account — an unreachable or missing bucket surfaces UnknownStoreException / NoSuchBucket in the History Server log. An empty History Server list is normal until a job completes. The S3A connector jars are downloaded from Maven Central at container start on every driver tier and the History Server.

Important Notes

  • No authentication anywhere by default — every Web UI and the cluster port are open to whoever can reach them. Keep publicAccess: false and reach UIs via cpln port-forward. Setting publicAccess: true puts the unauthenticated Master and History UIs on the public internet.
  • A driver must set spark.driver.host to its pod IP (see Submitting Jobs) — without it the driver advertises a non-routable pod hostname and the job stalls relaunching executors.
  • The History Server needs an existing bucket — enable historyServer.enabled and point storage.* at a real bucket and cloud account, or it crash-loops. Access is keyless via the workload identity.
  • Spark Connect clients need their own dependencies — the apache/spark image ships only the server side. Install pyspark[connect] (or pandas, pyarrow, and grpcio) in the client environment.
  • Switching provider (aws ↔ gcp) needs a fresh install, not an upgrade — cloud-binding blocks on an identity are never cleared on update.
  • internalAccess.type: workload-list must include the cluster’s own workloads — the master, workers, and Connect server address each other over the GVC network. same-gvc (the default) avoids this.
  • Worker scratch is ephemeral and workers.memory is Spark’s offer to executors, not the container limit — keep it below workers.maxMemory. Master loss is minutes of downtime, not data loss (there is no Master HA in this version).

External References

Spark Standalone Mode

How the standalone master and workers operate

Monitoring & History Server

Event logs, the History Server, and metrics

Spark Connect

Thin/remote client submission over gRPC

Submitting Applications

spark-submit options and application packaging

Hadoop-AWS S3A

The S3A connector used for event-log storage

Spark Template

View the source files, default values, and chart definition