Overview
Spark deploys an Apache Spark standalone cluster — a distributed engine for batch ETL, large-scale data transformation, and SQL/DataFrame analytics. The template provisions a master (cluster manager and Web UI), a horizontally-scalable worker tier, an optional Spark Connect gRPC server for thin/remote clients, and an optional History Server that reads completed applications’ event logs back from an S3 or GCS bucket.Architecture
- Master — cluster manager and Web UI on
:8080, cluster RPC on:7077. Fixed single replica. - Worker — the executor host tier, with its own Web UI on
:8081.workers.replicas(default1); more than one forms a multi-worker cluster. - Spark Connect (optional) — a gRPC server on
:15002for thin/remote client submission. Off by default. - History Server (optional) — a Web UI on
:18080that lists completed applications by reading their event logs back from object storage. Off by default; requires an S3/GCS bucket.
What Gets Created
- Standard Spark Master Workload —
{release}-spark-master, the cluster manager and Web UI. One replica; UI port:8080(http), RPC port:7077(tcp). - Standard Spark Worker Workload —
{release}-spark-worker, the executor host tier.workers.replicasreplicas; Web UI:8081and RPC:7078(tcp, internal only). - Standard Spark Connect Workload (optional) —
{release}-spark-connect, a gRPC server on:15002for remote client submission. Created only whenconnect.enabled: true. - Standard History Server Workload (optional) —
{release}-spark-history, the completed-application dashboard on:18080. Created only whenhistoryServer.enabled: true. - Config Secret —
{release}-spark-conf, an opaque secret (encoding: plain) holding the renderedspark-defaults.conf(reverse-proxy, pinned RPC ports, and event-log/S3A settings). Always created. - Identity & Policy — one shared identity with a policy granting
revealon the config secret. When the History Server is enabled, the identity also carries the cloud-account binding for keyless object-storage access.
This template does not create a GVC. You must deploy it into an existing GVC. No persistent volume is created either — worker scratch is ephemeral, and the History Server’s durable store is the object-storage bucket.
Prerequisites
- Default install: none. The master and a single worker come up with no external dependencies.
- History Server (
historyServer.enabled: true): an object-storage bucket and a Control Plane cloud account — AWS S3 or GCS. See Storage Setup.
UI
Browse, install, and manage templates visually
CLI
Manage templates from your terminal
Terraform
Declare templates in your Terraform configurations
Pulumi
Declare templates in your Pulumi programs
Configuration
The defaultvalues.yaml for this template:
Image
image— The Spark container image used by every tier. The chart is shipped and tested onapache/spark:4.0.4-scala2.13-java17-python3-ubuntu.
Resources
Each tier (master, workers, connect, historyServer) exposes its own minCpu / maxCpu / minMemory / maxMemory. maxCpu maps to the workload’s CPU limit and maxMemory to its memory limit, with minCpu / minMemory as the reservation.
Workers
workers.replicas— Number of worker replicas.1is the proven single shape; a value greater than1forms a multi-worker cluster and tasks are spread across replicas.workers.cores— Cores each worker offers to the cluster (SPARK_WORKER_CORES).workers.memory— Memory each worker offers to executors (SPARK_WORKER_MEMORY). This is Spark’s offer to executors, not the container limit — keep it belowworkers.maxMemory.
Spark Connect
connect.enabled— Whentrue, starts a Spark Connect gRPC server on:15002for thin/remote client submission. Off by default.
A PySpark Connect client needs its own dependencies. The
apache/spark image ships only the server side — install pyspark[connect] (or pandas, pyarrow, and grpcio) in the client environment.History Server
historyServer.enabled— Whentrue, starts the History Server on:18080, which lists completed applications by reading their event logs back from object storage. Off by default; it requiresstorage.*and a real bucket. See Storage Setup.
Object Storage
Configured only when the History Server is enabled. Access is keyless (UCI) — the workload identity federates with your cloud account and no access keys are stored in values.
The cluster’s driver tiers write event logs to this bucket and the History Server reads them back. See Storage Setup for the per-provider steps.
Extra Spark Configuration
sparkDefaults— A map ofkey: valueentries merged verbatim intospark-defaults.conf, for any Spark setting the template does not expose directly (for examplespark.sql.shuffle.partitions). See the Spark configuration reference.
Access
internalAccess.type— Controls which workloads may reach the cluster over the GVC network:
publicAccess.enabled— Whentrue, exposes the Master and History Server Web UIs on a public*.cpln.appcanonical endpoint. The worker and Connect internal UIs stay internal and are reached through the master’s reverse proxy. See Important Notes — the UIs are unauthenticated.
Submitting Jobs
Submit from the master container or any client workload in the GVC. A driver must advertise its pod IP — setspark.driver.host, or executors cannot connect back to the driver and the job stalls relaunching executors:
SparkPi run prints Pi is roughly 3.14.... For remote/thin-client submission without an exec shell, enable Spark Connect and point a client at sc://…:15002.
Connecting
All endpoints are internal GVC addresses. ReplaceRELEASE_NAME and GVC_NAME with your values.
The worker and running-application (
:4040) UIs are not exposed directly — the master serves them through its reverse proxy at /proxy/{id}/, so they are reachable through the master endpoint.
Reaching a Private UI
WithpublicAccess: false (the default), reach any Web UI in a browser with cpln port-forward:
http://localhost:8080. The History Server is reached the same way — cpln port-forward RELEASE_NAME-spark-history 18080:18080 --gvc GVC_NAME. There are no credentials: Spark’s Web UIs and cluster port are unauthenticated.
Storage Setup
The History Server reads completed applications’ event logs from a bucket that the cluster’s driver tiers also write to. Access is keyless — the workload identity federates with your cloud account, so no keys are stored in values. SethistoryServer.enabled: true and configure storage.*.
- AWS S3
- GCS
1
Create a bucket
Create your S3 bucket and set
storage.bucket and storage.region to match. Set storage.provider: aws.2
Set up a Cloud Account
If you do not have one, create a Cloud Account for your AWS account. Set
storage.cloudAccountName to its name.3
Create a bucket-scoped IAM policy
Create an AWS IAM policy with the JSON below (replace
YOUR_BUCKET_NAME), then set storage.policyName to the policy’s name:4
Set the prefix
Set
storage.prefix to the folder within the bucket where event logs are written.Important Notes
- No authentication anywhere by default — every Web UI and the cluster port are open to whoever can reach them. Keep
publicAccess: falseand reach UIs viacpln port-forward. SettingpublicAccess: trueputs the unauthenticated Master and History UIs on the public internet. - A driver must set
spark.driver.hostto its pod IP (see Submitting Jobs) — without it the driver advertises a non-routable pod hostname and the job stalls relaunching executors. - The History Server needs an existing bucket — enable
historyServer.enabledand pointstorage.*at a real bucket and cloud account, or it crash-loops. Access is keyless via the workload identity. - Spark Connect clients need their own dependencies — the
apache/sparkimage ships only the server side. Installpyspark[connect](orpandas,pyarrow, andgrpcio) in the client environment. - Switching provider (
aws↔gcp) needs a fresh install, not an upgrade — cloud-binding blocks on an identity are never cleared on update. internalAccess.type: workload-listmust include the cluster’s own workloads — the master, workers, and Connect server address each other over the GVC network.same-gvc(the default) avoids this.- Worker scratch is ephemeral and
workers.memoryis Spark’s offer to executors, not the container limit — keep it belowworkers.maxMemory. Master loss is minutes of downtime, not data loss (there is no Master HA in this version).
External References
Spark Standalone Mode
How the standalone master and workers operate
Monitoring & History Server
Event logs, the History Server, and metrics
Spark Connect
Thin/remote client submission over gRPC
Submitting Applications
spark-submit options and application packaging
Hadoop-AWS S3A
The S3A connector used for event-log storage
Spark Template
View the source files, default values, and chart definition