How to Set Up Docker Containers for Local Spark Development Environments

Clone the DataExpert-io/data-engineer-handbook repository and run make up in the intermediate-bootcamp/materials/3-spark-fundamentals directory to launch a complete local Spark environment with Iceberg, MinIO, and Jupyter Notebook accessible at http://localhost:8888.

The DataExpert-io/data-engineer-handbook repository provides a production-ready Docker composition for local Apache Spark development that eliminates cloud dependencies. This containerized stack bundles Spark with Apache Iceberg support, an S3-compatible object store, and a Jupyter interface into a single-command deployment, allowing data engineers to develop and test PySpark applications locally.

Architecture Overview

The Docker environment defined in intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml orchestrates four interconnected services across a private bridge network named iceberg_net:

  • spark-iceberg: The primary Spark container pre-configured with Iceberg support. It exposes port 8888 for Jupyter Notebook, port 8080 for the Spark UI, ports 10000/10001 for Thrift/SQL connections, and ports 4040-4042 for executor monitoring. This container mounts three host directories: ./warehouse to /home/iceberg/warehouse for the Iceberg warehouse, ./notebooks to /home/iceberg/notebooks/notebooks, and ./data to /home/iceberg/data for sample datasets.

  • rest: The Iceberg REST catalog service that bridges Spark queries to the underlying storage backend at s3://warehouse/.

  • minio: A local S3-compatible object store that persists Iceberg table data using dummy AWS credentials (admin/password).

  • mc: A helper container that initializes the MinIO bucket on startup, creating the warehouse bucket and applying public access policies before exiting.

All services communicate via the iceberg_net network, enabling DNS resolution by service name (e.g., minio, rest).

Prerequisites and Project Structure

Before launching the environment, ensure you have Docker Engine and Docker Compose installed locally. The Spark fundamentals module contains the following key files:

  • docker-compose.yaml: Defines the multi-service stack including volume mounts and environment variables.
  • Makefile: Provides convenience targets make up and make down to manage the lifecycle.
  • requirements.txt: Lists Python dependencies such as PySpark and Iceberg Python bindings for local script execution.
  • README.md: Documents the training module and entry points for the provided notebooks.

Step-by-Step Setup Instructions

Follow these commands to deploy the local Spark development environment:

  1. Clone the repository and navigate to the Spark fundamentals directory:

    git clone https://github.com/DataExpert-io/data-engineer-handbook.git
    cd data-engineer-handbook/intermediate-bootcamp/materials/3-spark-fundamentals
  2. Install Python dependencies (optional for local PySpark scripts):

    pip install -r requirements.txt
  3. Launch the Docker stack:

    Using the Makefile:

    make up

    Or using Docker Compose directly:

    docker compose up -d

    This command starts the spark-iceberg, rest, minio, and mc containers in detached mode. The mc container waits for MinIO to become healthy, then creates the warehouse bucket and sets appropriate permissions.

  4. Verify all services are running:

    docker ps

    You should see four containers: spark-iceberg, iceberg-rest, minio, and mc (the latter will show as Exited after completing its initialization tasks).

Accessing the Development Environment

Once the containers are running, open your browser and navigate to http://localhost:8888 to access the Jupyter Notebook interface running inside the spark-iceberg container.

The spark-iceberg container comes pre-loaded with example notebooks such as event_data_pyspark.ipynb located in the mounted ./notebooks directory. You can create new notebooks or modify existing ones, with all changes persisted to your local filesystem through the volume mount.

Managing the Environment

To stop the development environment and remove the containers, run:

make down

Or execute:

docker compose down

This command stops all services and removes the containers while preserving your data in the ./warehouse, ./notebooks, and ./data directories.

Summary

  • The DataExpert-io/data-engineer-handbook provides a complete Docker containers for local Spark development setup in the intermediate-bootcamp/materials/3-spark-fundamentals path.
  • The stack includes Spark with Iceberg support, an Iceberg REST catalog, MinIO for S3-compatible storage, and a Jupyter Notebook interface.
  • Launch the environment with make up or docker compose up -d, then access Jupyter at http://localhost:8888.
  • The docker-compose.yaml file configures persistent storage via host volume mounts for the warehouse, notebooks, and data directories.
  • The mc helper container automatically initializes the MinIO bucket on first startup, eliminating manual configuration.

Frequently Asked Questions

What ports need to be available on my host machine?

The spark-iceberg container requires ports 8888 (Jupyter), 8080 (Spark UI), 10000 and 10001 (Thrift/SQL), and 4040 through 4042 (Spark executor UI) to be free on your localhost. Ensure no other services are bound to these ports before running docker compose up.

How do I persist data between container restarts?

Data persists automatically through Docker volume mounts defined in the docker-compose.yaml file. The ./warehouse directory stores Iceberg table metadata and data, ./notebooks preserves your Jupyter files, and ./data maintains sample datasets. These local directories remain intact even after running docker compose down.

Can I use this setup without the Makefile?

Yes. While the Makefile in intermediate-bootcamp/materials/3-spark-fundamentals provides convenient make up and make down commands, you can interact directly with Docker Compose. Use docker compose up -d to start the stack and docker compose down to stop it. The Makefile simply wraps these commands for brevity.

What credentials should I use to connect to MinIO?

The environment uses dummy AWS credentials that are pre-configured in the container environment variables: set the access key to admin and the secret key to password. These credentials work for the MinIO endpoint at http://localhost:9000 and allow Spark to read and write Iceberg tables via the S3 API without requiring actual cloud infrastructure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →