> ## Documentation Index
> Fetch the complete documentation index at: https://docs.controlplane.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Spark

> Deploy an Apache Spark standalone cluster on Control Plane — a master, scalable workers, an optional Spark Connect gRPC server, and an optional S3/GCS-backed History Server for distributed batch ETL, SQL, and DataFrame analytics.

## Overview

Spark deploys an [Apache Spark](https://spark.apache.org/) standalone cluster — a distributed engine for batch ETL, large-scale data transformation, and SQL/DataFrame analytics. The template provisions a master (cluster manager and Web UI), a horizontally-scalable worker tier, an optional Spark Connect gRPC server for thin/remote clients, and an optional History Server that reads completed applications' event logs back from an S3 or GCS bucket.

### Architecture

* **Master** — cluster manager and Web UI on `:8080`, cluster RPC on `:7077`. Fixed single replica.
* **Worker** — the executor host tier, with its own Web UI on `:8081`. `workers.replicas` (default `1`); more than one forms a multi-worker cluster.
* **Spark Connect** *(optional)* — a gRPC server on `:15002` for thin/remote client submission. Off by default.
* **History Server** *(optional)* — a Web UI on `:18080` that lists completed applications by reading their event logs back from object storage. Off by default; requires an S3/GCS bucket.

### What Gets Created

* **Standard Spark Master Workload** — `{release}-spark-master`, the cluster manager and Web UI. One replica; UI port `:8080` (`http`), RPC port `:7077` (`tcp`).
* **Standard Spark Worker Workload** — `{release}-spark-worker`, the executor host tier. `workers.replicas` replicas; Web UI `:8081` and RPC `:7078` (`tcp`, internal only).
* **Standard Spark Connect Workload** *(optional)* — `{release}-spark-connect`, a gRPC server on `:15002` for remote client submission. Created only when `connect.enabled: true`.
* **Standard History Server Workload** *(optional)* — `{release}-spark-history`, the completed-application dashboard on `:18080`. Created only when `historyServer.enabled: true`.
* **Config Secret** — `{release}-spark-conf`, an [opaque secret](/guides/create-secret/opaque) (`encoding: plain`) holding the rendered `spark-defaults.conf` (reverse-proxy, pinned RPC ports, and event-log/S3A settings). Always created.
* **Identity & Policy** — one shared identity with a policy granting `reveal` on the config secret. When the History Server is enabled, the identity also carries the cloud-account binding for keyless object-storage access.

<Note>
  This template does not create a GVC. You must deploy it into an existing GVC. No persistent volume is created either — worker scratch is ephemeral, and the History Server's durable store is the object-storage bucket.
</Note>

## Prerequisites

* **Default install:** none. The master and a single worker come up with no external dependencies.
* **History Server (`historyServer.enabled: true`):** an object-storage bucket and a Control Plane [cloud account](https://docs.controlplane.com/guides/create-cloud-account) — AWS S3 or GCS. See [Storage Setup](#storage-setup).

Install the template using your preferred method:

<CardGroup cols={2}>
  <Card title="UI" href="/template-catalog/install-manage/ui" icon="laptop">
    Browse, install, and manage templates visually
  </Card>

  <Card title="CLI" href="/template-catalog/install-manage/cli" icon="terminal">
    Manage templates from your terminal
  </Card>

  <Card title="Terraform" href="/template-catalog/install-manage/terraform" icon={<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 128 128"><g fill-rule="evenodd"><path d="M77.941 44.5v36.836L46.324 62.918V26.082zm0 0" fill="#5c4ee5"/><path d="M81.41 81.336l31.633-18.418V26.082L81.41 44.5zm0 0" fill="#4040b2"/><path d="M11.242 42.36L42.86 60.776V23.941L11.242 5.523zm0 0M77.941 85.375L46.324 66.957v36.82l31.617 18.418zm0 0" fill="#5c4ee5"/></g></svg>}>
    Declare templates in your Terraform configurations
  </Card>

  <Card
    title="Pulumi"
    href="/template-catalog/install-manage/pulumi"
    icon={<svg xmlns="http://www.w3.org/2000/svg" fill="none" viewBox="0 0 24 24" id="Pulumi-Icon--Streamline-Svg-Logos" height="24" width="24">
    <desc>
        Pulumi Icon Streamline Icon: https://streamlinehq.com
    </desc>
    <path fill="#f26e7e" d="M4.683025 13.3318c0.869125 -0.5018 0.870575 -2.1264 0.003225 -3.62865s-2.27504 -2.313275 -3.1441725 -1.811475C0.672945 8.3935 0.6715 10.0181 1.53885 11.52035c0.86735 1.502275 2.27505 2.313275 3.144175 1.81145Zm0.0052 3.2167c0.86735 1.502275 0.865925 3.126875 -0.003225 3.628675 -0.86915 0.5018 -2.2768275 -0.309225 -3.144175 -1.81145 -0.8673525 -1.50225 -0.8659075 -3.126875 0.003225 -3.628675 0.8691325 -0.5018 2.276825 0.309225 3.144175 1.81145Zm5.922875 3.4243c0.86735 1.50225 0.8659 3.126775 -0.003225 3.62875 -0.869125 0.501775 -2.27685 -0.309325 -3.1442 -1.81155 -0.867325 -1.50225 -0.865875 -3.12685 0.00325 -3.628675 0.869125 -0.5018 2.276825 0.309225 3.144175 1.811475Zm-0.001925 -6.845275c0.86735 1.50225 0.8659 3.12685 -0.003225 3.628675 -0.869125 0.5018 -2.276825 -0.309225 -3.144175 -1.811475 -0.86735 -1.50225 -0.8659 -3.12685 0.003225 -3.62865 0.869125 -0.501825 2.276825 0.3092 3.144175 1.81145Z" stroke-width="0.25"></path>
    <path fill="#8a3391" d="M22.45775 11.524125c0.86725 -1.502225 0.865925 -3.12685 -0.003225 -3.62865 -0.869125 -0.501825 -2.276825 0.3092 -3.144175 1.811475 -0.86735 1.50225 -0.8659 3.126825 0.003225 3.62865 0.869125 0.501825 2.276825 -0.3092 3.144175 -1.811475Zm0.000175 3.2151c0.869075 0.5018 0.870625 2.1264 0.003225 3.62865 -0.86735 1.50225 -2.27505 2.313275 -3.144175 1.81145 -0.869125 -0.5018 -0.870575 -2.126425 -0.003225 -3.62865 0.86735 -1.50225 2.27505 -2.313275 3.144175 -1.81145ZM16.536225 18.157875c0.86915 0.501825 0.8706 2.126425 0.00325 3.628675 -0.86735 1.502125 -2.275075 2.313225 -3.1442 1.81145 -0.869125 -0.50175 -0.870575 -2.126425 -0.003225 -3.62865 0.867375 -1.502275 2.27505 -2.3133 3.144175 -1.811475Zm-0.003325 -6.843775c0.869125 0.5018 0.870575 2.126425 0.003225 3.628675s-2.27505 2.313275 -3.1442 1.811475c-0.869125 -0.501825 -0.870575 -2.126425 -0.003225 -3.628675 0.86735 -1.502275 2.27505 -2.313275 3.1442 -1.811475Z" stroke-width="0.25"></path>
    <path fill="#f7bf2a" d="M15.138225 2.06721c0 1.003615 -1.40625 1.817215 -3.14095 1.817215 -1.7347 0 -3.14095 -0.8136 -3.14095 -1.817215C8.856325 1.06359 10.262575 0.25 11.997275 0.25c1.7347 0 3.14095 0.81359 3.14095 1.81721ZM9.2166 5.482375c0 1.003625 -1.40625 1.8172 -3.14095 1.8172 -1.7347 0 -3.14095 -0.813575 -3.14095 -1.8172s1.40625 -1.817225 3.14095 -1.817225c1.7347 0 3.14095 0.8136 3.14095 1.817225Zm8.71005 1.8172c1.7347 0 3.14095 -0.813575 3.14095 -1.8172s-1.40625 -1.817225 -3.14095 -1.817225c-1.7347 0 -3.14095 0.8136 -3.14095 1.817225s1.40625 1.8172 3.14095 1.8172Zm-2.788425 1.605625c0 1.003625 -1.40625 1.8172 -3.14095 1.8172 -1.7347 0 -3.14095 -0.813575 -3.14095 -1.8172 0 -1.0036 1.40625 -1.8172 3.14095 -1.8172 1.7347 0 3.14095 0.8136 3.14095 1.8172Z" stroke-width="0.25"></path>
    </svg>}
  >
    Declare templates in your Pulumi programs
  </Card>
</CardGroup>

## Configuration

The default `values.yaml` for this template:

```yaml theme={null}
# ─── Image ───────────────────────────────────────────────────────
image: apache/spark:4.0.4-scala2.13-java17-python3-ubuntu

# ─── Master (cluster manager + Web UI :8080) ─────────────────────
master:
  minCpu: 250m
  maxCpu: 1
  minMemory: 512Mi
  maxMemory: 1Gi

# ─── Workers (executor host tier) ────────────────────────────────
workers:
  replicas: 1        # 1 = proven single shape; >1 forms a multi-worker cluster (see README)
  cores: 2           # SPARK_WORKER_CORES — cores this worker offers to the cluster
  memory: 2g         # SPARK_WORKER_MEMORY — memory offered to executors (keep < maxMemory)
  minCpu: 500m
  maxCpu: 2
  minMemory: 1Gi
  maxMemory: 4Gi

# ─── Spark Connect (optional remote submission, gRPC :15002) ─────
connect:
  enabled: false
  minCpu: 250m
  maxCpu: 1
  minMemory: 512Mi
  maxMemory: 1Gi

# ─── History Server (persistent completed-app dashboard, :18080) ──
# Requires object storage (bucket + cloud account) — see README "Storage setup".
historyServer:
  enabled: false
  minCpu: 250m
  maxCpu: 1
  minMemory: 512Mi
  maxMemory: 1Gi

# ─── Object storage (required when historyServer.enabled) ────────
# Keyless (UCI): create a cpln cloud account + a bucket-scoped IAM policy; no keys in values.
storage:
  provider: aws              # aws | gcp
  bucket: my-spark-events    # bucket you created for event logs
  prefix: spark/events       # folder within the bucket
  region: us-east-1          # aws only
  cloudAccountName: my-spark-cloud-account   # cpln cloud account linked to your provider
  policyName: my-spark-events-policy         # aws only: bucket-scoped IAM policy name

# ─── Extra Spark configuration (merged into spark-defaults.conf) ──
sparkDefaults: {}
  # spark.sql.shuffle.partitions: "200"

# ─── Access ──────────────────────────────────────────────────────
internalAccess:
  type: same-gvc      # same-gvc (recommended) | same-org | workload-list
  # workloads:        # only when type is workload-list
  #   - //gvc/GVC/workload/WORKLOAD

publicAccess:
  enabled: false      # true exposes the Master + History Web UIs publicly — NO AUTH; see README
```

### Image

* `image` — The Spark container image used by every tier. The chart is shipped and tested on `apache/spark:4.0.4-scala2.13-java17-python3-ubuntu`.

### Resources

Each tier (`master`, `workers`, `connect`, `historyServer`) exposes its own `minCpu` / `maxCpu` / `minMemory` / `maxMemory`. `maxCpu` maps to the workload's CPU limit and `maxMemory` to its memory limit, with `minCpu` / `minMemory` as the reservation.

### Workers

* `workers.replicas` — Number of worker replicas. `1` is the proven single shape; a value greater than `1` forms a multi-worker cluster and tasks are spread across replicas.
* `workers.cores` — Cores each worker offers to the cluster (`SPARK_WORKER_CORES`).
* `workers.memory` — Memory each worker offers to executors (`SPARK_WORKER_MEMORY`). This is Spark's *offer* to executors, not the container limit — keep it below `workers.maxMemory`.

### Spark Connect

* `connect.enabled` — When `true`, starts a [Spark Connect](https://spark.apache.org/docs/latest/spark-connect-overview.html) gRPC server on `:15002` for thin/remote client submission. Off by default.

<Note>
  A PySpark Connect client needs its own dependencies. The `apache/spark` image ships only the server side — install `pyspark[connect]` (or `pandas`, `pyarrow`, and `grpcio`) in the client environment.
</Note>

### History Server

* `historyServer.enabled` — When `true`, starts the History Server on `:18080`, which lists completed applications by reading their event logs back from object storage. Off by default; it **requires** `storage.*` and a real bucket. See [Storage Setup](#storage-setup).

### Object Storage

Configured only when the History Server is enabled. Access is keyless (UCI) — the workload identity federates with your cloud account and no access keys are stored in values.

| Field                      | Description                                              |
| -------------------------- | -------------------------------------------------------- |
| `storage.provider`         | `aws` or `gcp`.                                          |
| `storage.bucket`           | The bucket you created for event logs.                   |
| `storage.prefix`           | Folder within the bucket where event logs are written.   |
| `storage.region`           | Bucket region. **AWS only.**                             |
| `storage.cloudAccountName` | The Control Plane cloud account linked to your provider. |
| `storage.policyName`       | The bucket-scoped IAM policy name. **AWS only.**         |

The cluster's driver tiers write event logs to this bucket and the History Server reads them back. See [Storage Setup](#storage-setup) for the per-provider steps.

### Extra Spark Configuration

* `sparkDefaults` — A map of `key: value` entries merged verbatim into `spark-defaults.conf`, for any Spark setting the template does not expose directly (for example `spark.sql.shuffle.partitions`). See the [Spark configuration reference](https://spark.apache.org/docs/latest/configuration.html).

### Access

* `internalAccess.type` — Controls which workloads may reach the cluster over the GVC network:

| Type            | Description                                              |
| --------------- | -------------------------------------------------------- |
| `same-gvc`      | All workloads in the same GVC (recommended).             |
| `same-org`      | All workloads in the same organization.                  |
| `workload-list` | Only the workloads listed in `internalAccess.workloads`. |

<Warning>
  `internalAccess.type: workload-list` **must include the cluster's own workloads** — the master, workers, and Connect server address each other over the GVC network, so an incomplete list breaks the cluster itself. `same-gvc` (the default) avoids this.
</Warning>

* `publicAccess.enabled` — When `true`, exposes the **Master** and **History Server** Web UIs on a public `*.cpln.app` canonical endpoint. The worker and Connect internal UIs stay internal and are reached through the master's reverse proxy. See [Important Notes](#important-notes) — the UIs are unauthenticated.

## Submitting Jobs

Submit from the master container or any client workload in the GVC. **A driver must advertise its pod IP** — set `spark.driver.host`, or executors cannot connect back to the driver and the job stalls relaunching executors:

```bash theme={null}
cpln workload exec RELEASE_NAME-spark-master --gvc GVC_NAME --container spark-master -- bash -c '
  export SPARK_LOCAL_IP=$(hostname -i | awk "{print \$1}")
  /opt/spark/bin/spark-submit \
    --master spark://RELEASE_NAME-spark-master.GVC_NAME.cpln.local:7077 \
    --conf spark.driver.host=$SPARK_LOCAL_IP \
    --class org.apache.spark.examples.SparkPi \
    /opt/spark/examples/jars/spark-examples_2.13-4.0.4.jar 20'
```

A successful `SparkPi` run prints `Pi is roughly 3.14...`. For remote/thin-client submission without an exec shell, enable [Spark Connect](#spark-connect) and point a client at `sc://…:15002`.

## Connecting

All endpoints are internal GVC addresses. Replace `RELEASE_NAME` and `GVC_NAME` with your values.

| Target               | Endpoint                                                    | Notes                                             |
| -------------------- | ----------------------------------------------------------- | ------------------------------------------------- |
| Cluster RPC (submit) | `RELEASE_NAME-spark-master.GVC_NAME.cpln.local:7077`        | `spark://…:7077`; internal.                       |
| Master Web UI        | `RELEASE_NAME-spark-master.GVC_NAME.cpln.local:8080`        | Worker and application UIs are proxied behind it. |
| Spark Connect        | `sc://RELEASE_NAME-spark-connect.GVC_NAME.cpln.local:15002` | When `connect.enabled: true`.                     |
| History Server       | `RELEASE_NAME-spark-history.GVC_NAME.cpln.local:18080`      | When `historyServer.enabled: true`.               |

The worker and running-application (`:4040`) UIs are not exposed directly — the master serves them through its reverse proxy at `/proxy/{id}/`, so they are reachable through the master endpoint.

### Reaching a Private UI

With `publicAccess: false` (the default), reach any Web UI in a browser with [`cpln port-forward`](/cli-reference/commands/port-forward):

```bash theme={null}
cpln port-forward RELEASE_NAME-spark-master 8080:8080 --gvc GVC_NAME
```

Then open `http://localhost:8080`. The History Server is reached the same way — `cpln port-forward RELEASE_NAME-spark-history 18080:18080 --gvc GVC_NAME`. There are no credentials: Spark's Web UIs and cluster port are unauthenticated.

## Storage Setup

The History Server reads completed applications' event logs from a bucket that the cluster's driver tiers also write to. Access is keyless — the workload identity federates with your cloud account, so no keys are stored in values. Set `historyServer.enabled: true` and configure `storage.*`.

<Tabs>
  <Tab title="AWS S3">
    <Steps>
      <Step title="Create a bucket">
        Create your S3 bucket and set `storage.bucket` and `storage.region` to match. Set `storage.provider: aws`.
      </Step>

      <Step title="Set up a Cloud Account">
        If you do not have one, [create a Cloud Account](https://docs.controlplane.com/guides/create-cloud-account) for your AWS account. Set `storage.cloudAccountName` to its name.
      </Step>

      <Step title="Create a bucket-scoped IAM policy">
        Create an AWS IAM policy with the JSON below (replace `YOUR_BUCKET_NAME`), then set `storage.policyName` to the policy's name:

        ```json theme={null}
        {
            "Version": "2012-10-17",
            "Statement": [
                {
                    "Effect": "Allow",
                    "Action": [
                        "s3:GetObject",
                        "s3:PutObject",
                        "s3:DeleteObject",
                        "s3:ListBucket",
                        "s3:GetBucketLocation",
                        "s3:AbortMultipartUpload",
                        "s3:ListBucketMultipartUploads",
                        "s3:ListMultipartUploadParts"
                    ],
                    "Resource": [
                        "arn:aws:s3:::YOUR_BUCKET_NAME",
                        "arn:aws:s3:::YOUR_BUCKET_NAME/*"
                    ]
                }
            ]
        }
        ```
      </Step>

      <Step title="Set the prefix">
        Set `storage.prefix` to the folder within the bucket where event logs are written.
      </Step>
    </Steps>
  </Tab>

  <Tab title="GCS">
    <Steps>
      <Step title="Create a bucket">
        Create your GCS bucket and set `storage.bucket` to its name. Set `storage.provider: gcp`.
      </Step>

      <Step title="Set up a Cloud Account">
        [Create a Cloud Account](https://docs.controlplane.com/guides/create-cloud-account) for your GCP account. Set `storage.cloudAccountName` to its name.
      </Step>

      <Step title="Grant the Storage Admin role">
        Grant the GCP service account created for the Cloud Account the **Storage Admin** (`roles/storage.objectAdmin`) role on the bucket. Set `storage.prefix` to the folder for event logs. `region` and `policyName` are not used for GCP.
      </Step>
    </Steps>
  </Tab>
</Tabs>

<Warning>
  **The History Server crash-loops if the bucket does not exist.** Enable it only with a **real** bucket and cloud account — an unreachable or missing bucket surfaces `UnknownStoreException` / `NoSuchBucket` in the History Server log. An empty History Server list is normal until a job completes. The S3A connector jars are downloaded from Maven Central at container start on every driver tier and the History Server.
</Warning>

## Important Notes

* **No authentication anywhere by default** — every Web UI and the cluster port are open to whoever can reach them. Keep `publicAccess: false` and reach UIs via [`cpln port-forward`](#reaching-a-private-ui). Setting `publicAccess: true` puts the unauthenticated Master and History UIs on the public internet.
* **A driver must set `spark.driver.host` to its pod IP** (see [Submitting Jobs](#submitting-jobs)) — without it the driver advertises a non-routable pod hostname and the job stalls relaunching executors.
* **The History Server needs an existing bucket** — enable `historyServer.enabled` **and** point `storage.*` at a real bucket and cloud account, or it crash-loops. Access is keyless via the workload identity.
* **Spark Connect clients need their own dependencies** — the `apache/spark` image ships only the server side. Install `pyspark[connect]` (or `pandas`, `pyarrow`, and `grpcio`) in the client environment.
* **Switching provider (`aws` ↔ `gcp`) needs a fresh install**, not an upgrade — cloud-binding blocks on an identity are never cleared on update.
* **`internalAccess.type: workload-list` must include the cluster's own workloads** — the master, workers, and Connect server address each other over the GVC network. `same-gvc` (the default) avoids this.
* **Worker scratch is ephemeral** and `workers.memory` is Spark's offer to executors, not the container limit — keep it below `workers.maxMemory`. Master loss is minutes of downtime, not data loss (there is no Master HA in this version).

## External References

<CardGroup cols={2}>
  <Card title="Spark Standalone Mode" icon="server" href="https://spark.apache.org/docs/latest/spark-standalone.html">
    How the standalone master and workers operate
  </Card>

  <Card title="Monitoring & History Server" icon="chart-line" href="https://spark.apache.org/docs/latest/monitoring.html">
    Event logs, the History Server, and metrics
  </Card>

  <Card title="Spark Connect" icon="plug" href="https://spark.apache.org/docs/latest/spark-connect-overview.html">
    Thin/remote client submission over gRPC
  </Card>

  <Card title="Submitting Applications" icon="paper-plane" href="https://spark.apache.org/docs/latest/submitting-applications.html">
    spark-submit options and application packaging
  </Card>

  <Card title="Hadoop-AWS S3A" icon="aws" href="https://hadoop.apache.org/docs/r3.4.1/hadoop-aws/tools/hadoop-aws/index.html">
    The S3A connector used for event-log storage
  </Card>

  <Card title="Spark Template" icon="github" href="https://github.com/controlplane-com/templates/tree/main/spark">
    View the source files, default values, and chart definition
  </Card>
</CardGroup>
