How to Set Up Docker Containers for Local Spark Development Environments
Clone the DataExpert-io/data-engineer-handbook repository and run make up in the intermediate-bootcamp/materials/3-spark-fundamentals directory to launch a complete local Spark environment with Iceberg, MinIO, and Jupyter Notebook accessible at http://localhost:8888.
The DataExpert-io/data-engineer-handbook repository provides a production-ready Docker composition for local Apache Spark development that eliminates cloud dependencies. This containerized stack bundles Spark with Apache Iceberg support, an S3-compatible object store, and a Jupyter interface into a single-command deployment, allowing data engineers to develop and test PySpark applications locally.
Architecture Overview
The Docker environment defined in intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml orchestrates four interconnected services across a private bridge network named iceberg_net:
-
spark-iceberg: The primary Spark container pre-configured with Iceberg support. It exposes port
8888for Jupyter Notebook, port8080for the Spark UI, ports10000/10001for Thrift/SQL connections, and ports4040-4042for executor monitoring. This container mounts three host directories:./warehouseto/home/iceberg/warehousefor the Iceberg warehouse,./notebooksto/home/iceberg/notebooks/notebooks, and./datato/home/iceberg/datafor sample datasets. -
rest: The Iceberg REST catalog service that bridges Spark queries to the underlying storage backend at
s3://warehouse/. -
minio: A local S3-compatible object store that persists Iceberg table data using dummy AWS credentials (
admin/password). -
mc: A helper container that initializes the MinIO bucket on startup, creating the
warehousebucket and applying public access policies before exiting.
All services communicate via the iceberg_net network, enabling DNS resolution by service name (e.g., minio, rest).
Prerequisites and Project Structure
Before launching the environment, ensure you have Docker Engine and Docker Compose installed locally. The Spark fundamentals module contains the following key files:
docker-compose.yaml: Defines the multi-service stack including volume mounts and environment variables.Makefile: Provides convenience targetsmake upandmake downto manage the lifecycle.requirements.txt: Lists Python dependencies such as PySpark and Iceberg Python bindings for local script execution.README.md: Documents the training module and entry points for the provided notebooks.
Step-by-Step Setup Instructions
Follow these commands to deploy the local Spark development environment:
-
Clone the repository and navigate to the Spark fundamentals directory:
git clone https://github.com/DataExpert-io/data-engineer-handbook.git cd data-engineer-handbook/intermediate-bootcamp/materials/3-spark-fundamentals -
Install Python dependencies (optional for local PySpark scripts):
pip install -r requirements.txt -
Launch the Docker stack:
Using the Makefile:
make upOr using Docker Compose directly:
docker compose up -dThis command starts the
spark-iceberg,rest,minio, andmccontainers in detached mode. Themccontainer waits for MinIO to become healthy, then creates thewarehousebucket and sets appropriate permissions. -
Verify all services are running:
docker psYou should see four containers:
spark-iceberg,iceberg-rest,minio, andmc(the latter will show asExitedafter completing its initialization tasks).
Accessing the Development Environment
Once the containers are running, open your browser and navigate to http://localhost:8888 to access the Jupyter Notebook interface running inside the spark-iceberg container.
The spark-iceberg container comes pre-loaded with example notebooks such as event_data_pyspark.ipynb located in the mounted ./notebooks directory. You can create new notebooks or modify existing ones, with all changes persisted to your local filesystem through the volume mount.
Managing the Environment
To stop the development environment and remove the containers, run:
make down
Or execute:
docker compose down
This command stops all services and removes the containers while preserving your data in the ./warehouse, ./notebooks, and ./data directories.
Summary
- The DataExpert-io/data-engineer-handbook provides a complete Docker containers for local Spark development setup in the
intermediate-bootcamp/materials/3-spark-fundamentalspath. - The stack includes Spark with Iceberg support, an Iceberg REST catalog, MinIO for S3-compatible storage, and a Jupyter Notebook interface.
- Launch the environment with
make upordocker compose up -d, then access Jupyter athttp://localhost:8888. - The
docker-compose.yamlfile configures persistent storage via host volume mounts for the warehouse, notebooks, and data directories. - The
mchelper container automatically initializes the MinIO bucket on first startup, eliminating manual configuration.
Frequently Asked Questions
What ports need to be available on my host machine?
The spark-iceberg container requires ports 8888 (Jupyter), 8080 (Spark UI), 10000 and 10001 (Thrift/SQL), and 4040 through 4042 (Spark executor UI) to be free on your localhost. Ensure no other services are bound to these ports before running docker compose up.
How do I persist data between container restarts?
Data persists automatically through Docker volume mounts defined in the docker-compose.yaml file. The ./warehouse directory stores Iceberg table metadata and data, ./notebooks preserves your Jupyter files, and ./data maintains sample datasets. These local directories remain intact even after running docker compose down.
Can I use this setup without the Makefile?
Yes. While the Makefile in intermediate-bootcamp/materials/3-spark-fundamentals provides convenient make up and make down commands, you can interact directly with Docker Compose. Use docker compose up -d to start the stack and docker compose down to stop it. The Makefile simply wraps these commands for brevity.
What credentials should I use to connect to MinIO?
The environment uses dummy AWS credentials that are pre-configured in the container environment variables: set the access key to admin and the secret key to password. These credentials work for the MinIO endpoint at http://localhost:9000 and allow Spark to read and write Iceberg tables via the S3 API without requiring actual cloud infrastructure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →