How to Set Up and Configure Apache Iceberg for a Data Lakehouse Architecture

Deploy a complete Apache Iceberg lakehouse locally using Docker Compose with Spark, an Iceberg REST catalog, and MinIO as an S3-compatible object store from the DataExpert-io/data-engineer-handbook repository.

Apache Iceberg brings ACID transactions, schema evolution, and efficient query planning to data lake environments, enabling a true lakehouse architecture. The DataExpert-io/data-engineer-handbook repository provides a production-ready Docker Compose stack found in intermediate-bootcamp/materials/3-spark-fundamentals/ that demonstrates how to configure these components to work together seamlessly. This setup allows data engineers to experiment with Iceberg tables using standard SQL and Spark without managing complex cloud infrastructure.

Understanding the Iceberg Lakehouse Architecture

The lakehouse pattern implemented in the docker-compose.yaml file combines three essential services to create a functional Iceberg environment. Apache Spark serves as the execution engine for reading and writing table data, while the Iceberg REST Catalog maintains centralized metadata management via a REST API. MinIO provides an S3-compatible object store that persists both the actual data files and table metadata in the warehouse bucket.

Prerequisites and Repository Setup

Before launching the environment, ensure you have Docker and Docker Compose installed on your local machine. Clone the repository and navigate to the Spark Fundamentals directory where the configuration files reside.

git clone https://github.com/DataExpert-io/data-engineer-handbook.git
cd intermediate-bootcamp/materials/3-spark-fundamentals

The Makefile in this directory simplifies container orchestration with simple make up and make down commands that wrap the underlying Docker Compose operations.

Configuring the Docker Compose Environment

The docker-compose.yaml file located at intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml defines the complete service topology. Each container connects through the isolated iceberg_net network, ensuring secure internal communication between components with DNS aliases like warehouse.minio for service discovery.

Spark Configuration

The Spark service uses the tabulario/spark-iceberg image and exposes multiple ports including 8888 for Jupyter notebooks and 8080 for the Spark UI. Environment variables configure S3 access credentials that match the MinIO configuration:

services:
  spark-iceberg:
    image: tabulario/spark-iceberg
    container_name: spark-iceberg
    networks: [iceberg_net]
    depends_on: [rest, minio]
    volumes:
      - ./warehouse:/home/iceberg/warehouse
      - ./notebooks:/home/iceberg/notebooks/notebooks
      - ./data:/home/iceberg/data
    environment:
      - AWS_ACCESS_KEY_ID=admin
      - AWS_SECRET_ACCESS_KEY=password
      - AWS_REGION=us-east-1
    ports: [8888:8888, 8080:8080, 10000:10000, 10001:10001, "4040-4042:4040-4042"]

Iceberg REST Catalog Setup

The REST catalog service runs the tabulario/iceberg-rest image and acts as the central metadata service. It configures the S3FileIO implementation (org.apache.iceberg.aws.s3.S3FileIO) to read and write data to the MinIO bucket at s3://warehouse/. This service enables table metadata persistence across Spark sessions and provides the REST API endpoint that Spark uses to resolve table locations.

MinIO Object Storage

MinIO launches with the minio/minio image using the credentials admin/password, which the Spark and catalog services share through environment variables. The mc sidecar container automatically creates the warehouse bucket on startup, eliminating manual configuration steps and ensuring the storage layer is ready before Spark attempts to write data.

Creating and Querying Iceberg Tables

Once the stack is running via make up, access the Jupyter notebook at http://localhost:8888 to interact with Iceberg tables. The bucket-joins-in-iceberg.ipynb notebook demonstrates advanced table creation and query optimization techniques specific to the Iceberg format.

Table Creation with Bucket Partitioning

Iceberg supports sophisticated partitioning strategies that optimize query performance beyond simple Hive-style directories. The following Scala syntax creates a table with identity partitioning on completion_date and bucketing on match_id:

spark.sql(
  """
  CREATE TABLE IF NOT EXISTS bootcamp.matches_bucketed (
      match_id STRING,
      is_team_game BOOLEAN,
      playlist_id STRING,
      completion_date TIMESTAMP
    )
   USING iceberg
   PARTITIONED BY (completion_date, bucket(16, match_id));
  """
)

Query Optimization with Bucket Joins

Bucket partitioning enables efficient join operations by co-locating related data in the same file groups. To demonstrate this optimization, disable broadcast joins and run a bucket-aware query that leverages the bucket(16, match_id) partitioning:

// Disable automatic broadcast joins to force bucket-based optimization
spark.conf.set("spark.sql.autoBroadcastJoinThreshold", "-1")

// Execute bucket join that avoids shuffling data across the network
spark.sql(
  """
  SELECT *
  FROM bootcamp.match_details_bucketed mdb
  JOIN bootcamp.matches_bucketed md
    ON mdb.match_id = md.match_id
   AND md.completion_date = DATE('2016-01-01')
  """
).explain()

PySpark Configuration Alternative

For Python-based workflows, configure the SparkSession with Iceberg extensions and catalog settings to interact with the same underlying tables:

from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("IcebergDemo") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.iceberg.spark.SparkSessionCatalog") \
    .config("spark.sql.catalog.spark_catalog.type", "hadoop") \
    .config("spark.sql.catalog.spark_catalog.warehouse", "s3://warehouse/") \
    .config("spark.sql.extensions", "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions") \
    .getOrCreate()

# Create bucket-partitioned table

spark.sql("""
CREATE TABLE IF NOT EXISTS bootcamp.matches_bucketed (
    match_id STRING,
    is_team_game BOOLEAN,
    playlist_id STRING,
    completion_date TIMESTAMP
) USING iceberg
PARTITIONED BY (completion_date, bucket(16, match_id))
""")

Summary

  • The DataExpert-io/data-engineer-handbook repository provides a complete Docker Compose stack for Apache Iceberg lakehouse experimentation with a single make up command.
  • The architecture combines Spark, Iceberg REST Catalog, and MinIO to deliver ACID transactions and schema evolution capabilities using the tabulario/spark-iceberg and tabulario/iceberg-rest images.
  • Configuration in docker-compose.yaml uses the S3FileIO implementation to persist data in the MinIO warehouse bucket with credentials shared across services.
  • Bucket partitioning with bucket(16, match_id) optimizes join performance by distributing data across 16 buckets based on hash values.
  • Access the demonstration notebook at localhost:8888 to run the bucket-joins-in-iceberg.ipynb examples and validate the lakehouse setup.

Frequently Asked Questions

What is the role of the Iceberg REST Catalog in this architecture?

The Iceberg REST Catalog serves as the centralized metadata service that exposes table information via a REST API. According to the DataExpert-io/data-engineer-handbook source code, it runs the tabulario/iceberg-rest image and configures the S3FileIO implementation to persist metadata in the MinIO object store, enabling multiple Spark sessions to share table state consistently through the rest service defined in docker-compose.yaml.

Why does the setup use MinIO instead of Amazon S3?

MinIO provides an S3-compatible object storage layer that runs locally without cloud credentials or costs. The docker-compose.yaml configures MinIO with the minio/minio image and automatically creates the warehouse bucket via the mc sidecar container, making it ideal for development and testing Iceberg lakehouse patterns while maintaining full API compatibility with AWS S3.

How does bucket partitioning improve query performance in Iceberg?

Bucket partitioning distributes data files across a fixed number of buckets based on a hash of the bucket column, in this case bucket(16, match_id). When joining tables on the bucketed column, Spark can avoid shuffling data across the network by matching bucket indices directly, as demonstrated in the bucket-joins-in-iceberg.ipynb notebook where spark.sql.autoBroadcastJoinThreshold is set to -1 to force this optimization and reveal the physical execution plan.

Can I use PySpark instead of Scala with this Iceberg setup?

Yes, the Docker Compose environment supports both languages interchangeably. The PySpark configuration requires setting the spark.sql.extensions to org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions and configuring the spark_catalog to use the Hadoop catalog type pointing to s3://warehouse/, allowing you to execute identical CREATE TABLE and SELECT statements using Python syntax against the same Iceberg tables.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →