# How to Set Up Apache Spark for Data Processing: A Complete Docker-Based Guide

> Master Apache Spark data processing with our Docker Compose guide. Set up a complete development environment including Spark, Iceberg, MinIO, and JupyterLab locally.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-08

---

**Apache Spark in the DataExpert handbook is deployed via Docker Compose as a complete development environment bundling Spark, Iceberg, MinIO, and a REST catalog, accessible through JupyterLab at localhost:8888.**

Setting up Apache Spark for data processing can be complex, but the DataExpert-io/data-engineer-handbook repository provides a containerized solution that eliminates configuration headaches. This setup uses Docker Compose to orchestrate a multi-service stack including Spark, Iceberg table format, and an S3-compatible object store. Whether you are learning PySpark fundamentals or building production-grade data pipelines, this environment provides everything needed to read, transform, and write large datasets.

## Architecture of the Spark Environment

The [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) file in `intermediate-bootcamp/materials/3-spark-fundamentals/` defines four interconnected services that work together to create a fully functional data lakehouse.

### Core Services

- **spark-iceberg**: Runs the `tabulario/spark-iceberg` image containing Spark 3.x with Iceberg support. This container mounts the local `data` directory and exposes the Spark UI on ports **4040-4042**, the Thrift server on **10000-10001**, and JupyterLab on **8888**.
- **rest**: Hosts the Iceberg REST catalog (`tabulario/iceberg-rest`) that Spark queries to resolve table metadata and schema information.
- **minio**: Provides an S3-compatible object store that persists the Iceberg warehouse data.
- **mc**: A helper container that automatically creates the `warehouse` bucket in MinIO and sets its access policy to public during startup.

All services communicate over the dedicated `iceberg_net` Docker network, allowing Spark to reference the catalog and storage using internal hostnames (`rest` and `minio`) rather than IP addresses.

## Prerequisites and Installation

Before launching the environment, install the Python dependencies required to run the notebooks and test suites.

```bash
pip install -r requirements.txt

```

The [`requirements.txt`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/requirements.txt) file includes **PySpark**, **chispa** (for DataFrame testing), and **pytest**. These packages allow you to execute Spark jobs locally and validate results using the provided test suites in `src/tests/`.

## Starting the Apache Spark Environment

Once dependencies are installed, initialize the entire stack using the convenience Makefile target or Docker Compose directly.

```bash
make up

# Or manually:

docker compose up -d

```

This command pulls the necessary images and starts all four containers in detached mode. Verify the services are healthy by checking container logs or navigating to the Spark UI at `http://localhost:4040`.

## Configuration and Environment Variables

The Spark container is pre-configured to use Iceberg as the default catalog. Critical environment variables injected via [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) include:

- `AWS_ACCESS_KEY_ID=admin` and `AWS_SECRET_ACCESS_KEY=password`: Dummy S3 credentials that authenticate Spark to the MinIO object store.
- `SPARK_SQL_CATALOG_SPARK_CATALOG=org.apache.iceberg.spark.SparkCatalog`: Sets Iceberg as the default table catalog.
- `SPARK_SQL_CATALOG_SPARK_CATALOG_WAREHOUSE=s3://warehouse/`: Points the catalog to the MinIO bucket created by the `mc` helper container.

These configurations allow seamless reading and writing of Iceberg tables without manual AWS setup.

## Running Your First PySpark Job

After starting the environment, open `http://localhost:8888` to access JupyterLab. The repository includes `notebooks/event_data_pyspark.ipynb`, which demonstrates reading CSV files from the mounted `data` directory and writing to Iceberg tables.

For programmatic job execution, reference the pattern used in [`src/jobs/monthly_user_site_hits_job.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/src/jobs/monthly_user_site_hits_job.py):

```python
from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("MonthlyUserSiteHits") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.iceberg.spark.SparkCatalog") \
    .config("spark.sql.catalog.spark_catalog.warehouse", "s3://warehouse/") \
    .getOrCreate()

df = spark.read.csv("data/events.csv", header=True, inferSchema=True)
df.createOrReplaceTempView("events")
result = spark.sql("""
    SELECT user_id, month, COUNT(*) AS hits
    FROM events
    GROUP BY user_id, month
""")
result.show()

```

This script initializes a Spark session with Iceberg catalog support, loads event data, and performs aggregations using Spark SQL.

## Testing Your Spark Jobs

Validate your data transformations using the pytest suites provided in `src/tests/`. These tests utilize **chispa** to compare DataFrame outputs against expected schemas and values.

```bash
python -m pytest

```

Running this command from the repository root executes all test files, ensuring your Spark logic handles edge cases and data types correctly before deployment.

## Troubleshooting Memory Issues

If you encounter `OutOfMemoryError: Java heap space` during large dataset processing, increase the JVM heap size allocated to the Spark driver. Modify the `spark-iceberg` service definition in [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) to include:

```yaml
environment:
  - SPARK_DRIVER_MEMORY=4g

```

Alternatively, adjust Docker Desktop resource limits to allocate more RAM to the container runtime.

## Summary

- **Apache Spark** in this repository runs as a Dockerized stack with Iceberg, MinIO, and a REST catalog defined in [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml).
- Access the development environment via **JupyterLab** at `http://localhost:8888` after running `make up`.
- Python dependencies in [`requirements.txt`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/requirements.txt) include PySpark and testing libraries like chispa.
- The `spark-iceberg` container uses dummy S3 credentials to read/write from the MinIO `warehouse` bucket at `s3://warehouse/`.
- Example implementations in [`src/jobs/monthly_user_site_hits_job.py`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/src/jobs/monthly_user_site_hits_job.py) demonstrate proper SparkSession configuration for Iceberg tables.
- Resolve memory errors by increasing `SPARK_DRIVER_MEMORY` or Docker resource limits.

## Frequently Asked Questions

### How do I connect to the Spark UI?

Navigate to `http://localhost:4040` in your browser while the containers are running. The Spark UI displays active jobs, stages, and executor metrics. If port 4040 is occupied, Spark automatically binds to 4041 or 4042 as configured in the [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) port range mapping.

### Can I use this setup for production workloads?

This configuration is optimized for local development and education. While the Iceberg REST catalog and MinIO storage provide production-like semantics, you should replace the dummy credentials (`admin`/`password`) with proper IAM roles, enable TLS encryption, and configure high-availability Spark masters before deploying to production.

### Where is the data actually stored?

The Iceberg warehouse data persists in the MinIO container at `s3://warehouse/`. Because MinIO stores objects on the container filesystem, data survives container restarts but not image rebuilds unless you mount a Docker volume. The `mc` helper container automatically creates this bucket and sets public access policies during the initial `docker compose up` execution.

### How do I add custom Python packages to the Spark environment?

Install additional packages locally via `pip install -r requirements.txt` so your IDE recognizes them. For packages needed inside the Spark container (like specific JDBC drivers), extend the `tabulario/spark-iceberg` image in a custom Dockerfile or mount JAR files into the container's Spark classpath directory.