How to Set Up Apache Spark for Data Processing: A Complete Docker-Based Guide

Apache Spark in the DataExpert handbook is deployed via Docker Compose as a complete development environment bundling Spark, Iceberg, MinIO, and a REST catalog, accessible through JupyterLab at localhost:8888.

Setting up Apache Spark for data processing can be complex, but the DataExpert-io/data-engineer-handbook repository provides a containerized solution that eliminates configuration headaches. This setup uses Docker Compose to orchestrate a multi-service stack including Spark, Iceberg table format, and an S3-compatible object store. Whether you are learning PySpark fundamentals or building production-grade data pipelines, this environment provides everything needed to read, transform, and write large datasets.

Architecture of the Spark Environment

The docker-compose.yaml file in intermediate-bootcamp/materials/3-spark-fundamentals/ defines four interconnected services that work together to create a fully functional data lakehouse.

Core Services

  • spark-iceberg: Runs the tabulario/spark-iceberg image containing Spark 3.x with Iceberg support. This container mounts the local data directory and exposes the Spark UI on ports 4040-4042, the Thrift server on 10000-10001, and JupyterLab on 8888.
  • rest: Hosts the Iceberg REST catalog (tabulario/iceberg-rest) that Spark queries to resolve table metadata and schema information.
  • minio: Provides an S3-compatible object store that persists the Iceberg warehouse data.
  • mc: A helper container that automatically creates the warehouse bucket in MinIO and sets its access policy to public during startup.

All services communicate over the dedicated iceberg_net Docker network, allowing Spark to reference the catalog and storage using internal hostnames (rest and minio) rather than IP addresses.

Prerequisites and Installation

Before launching the environment, install the Python dependencies required to run the notebooks and test suites.

pip install -r requirements.txt

The requirements.txt file includes PySpark, chispa (for DataFrame testing), and pytest. These packages allow you to execute Spark jobs locally and validate results using the provided test suites in src/tests/.

Starting the Apache Spark Environment

Once dependencies are installed, initialize the entire stack using the convenience Makefile target or Docker Compose directly.

make up

# Or manually:

docker compose up -d

This command pulls the necessary images and starts all four containers in detached mode. Verify the services are healthy by checking container logs or navigating to the Spark UI at http://localhost:4040.

Configuration and Environment Variables

The Spark container is pre-configured to use Iceberg as the default catalog. Critical environment variables injected via docker-compose.yaml include:

  • AWS_ACCESS_KEY_ID=admin and AWS_SECRET_ACCESS_KEY=password: Dummy S3 credentials that authenticate Spark to the MinIO object store.
  • SPARK_SQL_CATALOG_SPARK_CATALOG=org.apache.iceberg.spark.SparkCatalog: Sets Iceberg as the default table catalog.
  • SPARK_SQL_CATALOG_SPARK_CATALOG_WAREHOUSE=s3://warehouse/: Points the catalog to the MinIO bucket created by the mc helper container.

These configurations allow seamless reading and writing of Iceberg tables without manual AWS setup.

Running Your First PySpark Job

After starting the environment, open http://localhost:8888 to access JupyterLab. The repository includes notebooks/event_data_pyspark.ipynb, which demonstrates reading CSV files from the mounted data directory and writing to Iceberg tables.

For programmatic job execution, reference the pattern used in src/jobs/monthly_user_site_hits_job.py:

from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("MonthlyUserSiteHits") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.iceberg.spark.SparkCatalog") \
    .config("spark.sql.catalog.spark_catalog.warehouse", "s3://warehouse/") \
    .getOrCreate()

df = spark.read.csv("data/events.csv", header=True, inferSchema=True)
df.createOrReplaceTempView("events")
result = spark.sql("""
    SELECT user_id, month, COUNT(*) AS hits
    FROM events
    GROUP BY user_id, month
""")
result.show()

This script initializes a Spark session with Iceberg catalog support, loads event data, and performs aggregations using Spark SQL.

Testing Your Spark Jobs

Validate your data transformations using the pytest suites provided in src/tests/. These tests utilize chispa to compare DataFrame outputs against expected schemas and values.

python -m pytest

Running this command from the repository root executes all test files, ensuring your Spark logic handles edge cases and data types correctly before deployment.

Troubleshooting Memory Issues

If you encounter OutOfMemoryError: Java heap space during large dataset processing, increase the JVM heap size allocated to the Spark driver. Modify the spark-iceberg service definition in docker-compose.yaml to include:

environment:
  - SPARK_DRIVER_MEMORY=4g

Alternatively, adjust Docker Desktop resource limits to allocate more RAM to the container runtime.

Summary

  • Apache Spark in this repository runs as a Dockerized stack with Iceberg, MinIO, and a REST catalog defined in docker-compose.yaml.
  • Access the development environment via JupyterLab at http://localhost:8888 after running make up.
  • Python dependencies in requirements.txt include PySpark and testing libraries like chispa.
  • The spark-iceberg container uses dummy S3 credentials to read/write from the MinIO warehouse bucket at s3://warehouse/.
  • Example implementations in src/jobs/monthly_user_site_hits_job.py demonstrate proper SparkSession configuration for Iceberg tables.
  • Resolve memory errors by increasing SPARK_DRIVER_MEMORY or Docker resource limits.

Frequently Asked Questions

How do I connect to the Spark UI?

Navigate to http://localhost:4040 in your browser while the containers are running. The Spark UI displays active jobs, stages, and executor metrics. If port 4040 is occupied, Spark automatically binds to 4041 or 4042 as configured in the docker-compose.yaml port range mapping.

Can I use this setup for production workloads?

This configuration is optimized for local development and education. While the Iceberg REST catalog and MinIO storage provide production-like semantics, you should replace the dummy credentials (admin/password) with proper IAM roles, enable TLS encryption, and configure high-availability Spark masters before deploying to production.

Where is the data actually stored?

The Iceberg warehouse data persists in the MinIO container at s3://warehouse/. Because MinIO stores objects on the container filesystem, data survives container restarts but not image rebuilds unless you mount a Docker volume. The mc helper container automatically creates this bucket and sets public access policies during the initial docker compose up execution.

How do I add custom Python packages to the Spark environment?

Install additional packages locally via pip install -r requirements.txt so your IDE recognizes them. For packages needed inside the Spark container (like specific JDBC drivers), extend the tabulario/spark-iceberg image in a custom Dockerfile or mount JAR files into the container's Spark classpath directory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →