# How to Set Up Docker Containers for Local Spark Development Environments

> Set up local Spark development environments with Docker. Clone the DataExpert-io repository and use `make up` for a complete setup including Iceberg, MinIO, and Jupyter Notebook.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-12

---

**Clone the DataExpert-io/data-engineer-handbook repository and run `make up` in the `intermediate-bootcamp/materials/3-spark-fundamentals` directory to launch a complete local Spark environment with Iceberg, MinIO, and Jupyter Notebook accessible at `http://localhost:8888`.**

The DataExpert-io/data-engineer-handbook repository provides a production-ready Docker composition for local Apache Spark development that eliminates cloud dependencies. This containerized stack bundles Spark with Apache Iceberg support, an S3-compatible object store, and a Jupyter interface into a single-command deployment, allowing data engineers to develop and test PySpark applications locally.

## Architecture Overview

The Docker environment defined in [`intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml) orchestrates four interconnected services across a private bridge network named `iceberg_net`:

- **spark-iceberg**: The primary Spark container pre-configured with Iceberg support. It exposes port `8888` for Jupyter Notebook, port `8080` for the Spark UI, ports `10000/10001` for Thrift/SQL connections, and ports `4040-4042` for executor monitoring. This container mounts three host directories: `./warehouse` to `/home/iceberg/warehouse` for the Iceberg warehouse, `./notebooks` to `/home/iceberg/notebooks/notebooks`, and `./data` to `/home/iceberg/data` for sample datasets.

- **rest**: The Iceberg REST catalog service that bridges Spark queries to the underlying storage backend at `s3://warehouse/`.

- **minio**: A local S3-compatible object store that persists Iceberg table data using dummy AWS credentials (`admin`/`password`).

- **mc**: A helper container that initializes the MinIO bucket on startup, creating the `warehouse` bucket and applying public access policies before exiting.

All services communicate via the `iceberg_net` network, enabling DNS resolution by service name (e.g., `minio`, `rest`).

## Prerequisites and Project Structure

Before launching the environment, ensure you have Docker Engine and Docker Compose installed locally. The Spark fundamentals module contains the following key files:

- [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml): Defines the multi-service stack including volume mounts and environment variables.
- `Makefile`: Provides convenience targets `make up` and `make down` to manage the lifecycle.
- [`requirements.txt`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/requirements.txt): Lists Python dependencies such as PySpark and Iceberg Python bindings for local script execution.
- [`README.md`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/README.md): Documents the training module and entry points for the provided notebooks.

## Step-by-Step Setup Instructions

Follow these commands to deploy the local Spark development environment:

1. **Clone the repository and navigate to the Spark fundamentals directory:**

   ```bash
   git clone https://github.com/DataExpert-io/data-engineer-handbook.git
   cd data-engineer-handbook/intermediate-bootcamp/materials/3-spark-fundamentals
   ```

2. **Install Python dependencies (optional for local PySpark scripts):**

   ```bash
   pip install -r requirements.txt
   ```

3. **Launch the Docker stack:**

   Using the Makefile:
   ```bash
   make up
   ```

   
   Or using Docker Compose directly:
   ```bash
   docker compose up -d
   ```

   This command starts the `spark-iceberg`, `rest`, `minio`, and `mc` containers in detached mode. The `mc` container waits for MinIO to become healthy, then creates the `warehouse` bucket and sets appropriate permissions.

4. **Verify all services are running:**

   ```bash
   docker ps
   ```

   
   You should see four containers: `spark-iceberg`, `iceberg-rest`, `minio`, and `mc` (the latter will show as `Exited` after completing its initialization tasks).

## Accessing the Development Environment

Once the containers are running, open your browser and navigate to `http://localhost:8888` to access the Jupyter Notebook interface running inside the `spark-iceberg` container.

The `spark-iceberg` container comes pre-loaded with example notebooks such as `event_data_pyspark.ipynb` located in the mounted `./notebooks` directory. You can create new notebooks or modify existing ones, with all changes persisted to your local filesystem through the volume mount.

## Managing the Environment

To stop the development environment and remove the containers, run:

```bash
make down

```

Or execute:

```bash
docker compose down

```

This command stops all services and removes the containers while preserving your data in the `./warehouse`, `./notebooks`, and `./data` directories.

## Summary

- The **DataExpert-io/data-engineer-handbook** provides a complete **Docker containers for local Spark development** setup in the `intermediate-bootcamp/materials/3-spark-fundamentals` path.
- The stack includes **Spark with Iceberg support**, an **Iceberg REST catalog**, **MinIO** for S3-compatible storage, and a **Jupyter Notebook** interface.
- Launch the environment with **`make up`** or **`docker compose up -d`**, then access Jupyter at **`http://localhost:8888`**.
- The [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) file configures persistent storage via host volume mounts for the warehouse, notebooks, and data directories.
- The **`mc`** helper container automatically initializes the MinIO bucket on first startup, eliminating manual configuration.

## Frequently Asked Questions

### What ports need to be available on my host machine?

The `spark-iceberg` container requires ports `8888` (Jupyter), `8080` (Spark UI), `10000` and `10001` (Thrift/SQL), and `4040` through `4042` (Spark executor UI) to be free on your localhost. Ensure no other services are bound to these ports before running `docker compose up`.

### How do I persist data between container restarts?

Data persists automatically through Docker volume mounts defined in the [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) file. The `./warehouse` directory stores Iceberg table metadata and data, `./notebooks` preserves your Jupyter files, and `./data` maintains sample datasets. These local directories remain intact even after running `docker compose down`.

### Can I use this setup without the Makefile?

Yes. While the `Makefile` in `intermediate-bootcamp/materials/3-spark-fundamentals` provides convenient `make up` and `make down` commands, you can interact directly with Docker Compose. Use `docker compose up -d` to start the stack and `docker compose down` to stop it. The Makefile simply wraps these commands for brevity.

### What credentials should I use to connect to MinIO?

The environment uses dummy AWS credentials that are pre-configured in the container environment variables: set the access key to `admin` and the secret key to `password`. These credentials work for the MinIO endpoint at `http://localhost:9000` and allow Spark to read and write Iceberg tables via the S3 API without requiring actual cloud infrastructure.