# How to Set Up and Configure Apache Iceberg for a Data Lakehouse Architecture

> Set up Apache Iceberg for your data lakehouse locally with Docker Compose Spark a REST catalog and MinIO Learn to configure this powerful table format for efficient data management.

- Repository: [DataExpert.io/data-engineer-handbook](https://github.com/DataExpert-io/data-engineer-handbook)
- Tags: how-to-guide
- Published: 2026-08-12

---

**Deploy a complete Apache Iceberg lakehouse locally using Docker Compose with Spark, an Iceberg REST catalog, and MinIO as an S3-compatible object store from the DataExpert-io/data-engineer-handbook repository.**

Apache Iceberg brings ACID transactions, schema evolution, and efficient query planning to data lake environments, enabling a true lakehouse architecture. The DataExpert-io/data-engineer-handbook repository provides a production-ready Docker Compose stack found in `intermediate-bootcamp/materials/3-spark-fundamentals/` that demonstrates how to configure these components to work together seamlessly. This setup allows data engineers to experiment with Iceberg tables using standard SQL and Spark without managing complex cloud infrastructure.

## Understanding the Iceberg Lakehouse Architecture

The lakehouse pattern implemented in the [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) file combines three essential services to create a functional Iceberg environment. **Apache Spark** serves as the execution engine for reading and writing table data, while the **Iceberg REST Catalog** maintains centralized metadata management via a REST API. **MinIO** provides an S3-compatible object store that persists both the actual data files and table metadata in the `warehouse` bucket.

## Prerequisites and Repository Setup

Before launching the environment, ensure you have Docker and Docker Compose installed on your local machine. Clone the repository and navigate to the Spark Fundamentals directory where the configuration files reside.

```bash
git clone https://github.com/DataExpert-io/data-engineer-handbook.git
cd intermediate-bootcamp/materials/3-spark-fundamentals

```

The `Makefile` in this directory simplifies container orchestration with simple `make up` and `make down` commands that wrap the underlying Docker Compose operations.

## Configuring the Docker Compose Environment

The [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) file located at [`intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/intermediate-bootcamp/materials/3-spark-fundamentals/docker-compose.yaml) defines the complete service topology. Each container connects through the isolated `iceberg_net` network, ensuring secure internal communication between components with DNS aliases like `warehouse.minio` for service discovery.

### Spark Configuration

The Spark service uses the `tabulario/spark-iceberg` image and exposes multiple ports including `8888` for Jupyter notebooks and `8080` for the Spark UI. Environment variables configure S3 access credentials that match the MinIO configuration:

```yaml
services:
  spark-iceberg:
    image: tabulario/spark-iceberg
    container_name: spark-iceberg
    networks: [iceberg_net]
    depends_on: [rest, minio]
    volumes:
      - ./warehouse:/home/iceberg/warehouse
      - ./notebooks:/home/iceberg/notebooks/notebooks
      - ./data:/home/iceberg/data
    environment:
      - AWS_ACCESS_KEY_ID=admin
      - AWS_SECRET_ACCESS_KEY=password
      - AWS_REGION=us-east-1
    ports: [8888:8888, 8080:8080, 10000:10000, 10001:10001, "4040-4042:4040-4042"]

```

### Iceberg REST Catalog Setup

The REST catalog service runs the `tabulario/iceberg-rest` image and acts as the central metadata service. It configures the **S3FileIO** implementation (`org.apache.iceberg.aws.s3.S3FileIO`) to read and write data to the MinIO bucket at `s3://warehouse/`. This service enables table metadata persistence across Spark sessions and provides the REST API endpoint that Spark uses to resolve table locations.

### MinIO Object Storage

MinIO launches with the `minio/minio` image using the credentials `admin`/`password`, which the Spark and catalog services share through environment variables. The `mc` sidecar container automatically creates the `warehouse` bucket on startup, eliminating manual configuration steps and ensuring the storage layer is ready before Spark attempts to write data.

## Creating and Querying Iceberg Tables

Once the stack is running via `make up`, access the Jupyter notebook at `http://localhost:8888` to interact with Iceberg tables. The `bucket-joins-in-iceberg.ipynb` notebook demonstrates advanced table creation and query optimization techniques specific to the Iceberg format.

### Table Creation with Bucket Partitioning

Iceberg supports sophisticated partitioning strategies that optimize query performance beyond simple Hive-style directories. The following Scala syntax creates a table with identity partitioning on `completion_date` and bucketing on `match_id`:

```scala
spark.sql(
  """
  CREATE TABLE IF NOT EXISTS bootcamp.matches_bucketed (
      match_id STRING,
      is_team_game BOOLEAN,
      playlist_id STRING,
      completion_date TIMESTAMP
    )
   USING iceberg
   PARTITIONED BY (completion_date, bucket(16, match_id));
  """
)

```

### Query Optimization with Bucket Joins

Bucket partitioning enables efficient join operations by co-locating related data in the same file groups. To demonstrate this optimization, disable broadcast joins and run a bucket-aware query that leverages the `bucket(16, match_id)` partitioning:

```scala
// Disable automatic broadcast joins to force bucket-based optimization
spark.conf.set("spark.sql.autoBroadcastJoinThreshold", "-1")

// Execute bucket join that avoids shuffling data across the network
spark.sql(
  """
  SELECT *
  FROM bootcamp.match_details_bucketed mdb
  JOIN bootcamp.matches_bucketed md
    ON mdb.match_id = md.match_id
   AND md.completion_date = DATE('2016-01-01')
  """
).explain()

```

## PySpark Configuration Alternative

For Python-based workflows, configure the SparkSession with Iceberg extensions and catalog settings to interact with the same underlying tables:

```python
from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("IcebergDemo") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.iceberg.spark.SparkSessionCatalog") \
    .config("spark.sql.catalog.spark_catalog.type", "hadoop") \
    .config("spark.sql.catalog.spark_catalog.warehouse", "s3://warehouse/") \
    .config("spark.sql.extensions", "org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions") \
    .getOrCreate()

# Create bucket-partitioned table

spark.sql("""
CREATE TABLE IF NOT EXISTS bootcamp.matches_bucketed (
    match_id STRING,
    is_team_game BOOLEAN,
    playlist_id STRING,
    completion_date TIMESTAMP
) USING iceberg
PARTITIONED BY (completion_date, bucket(16, match_id))
""")

```

## Summary

- The **DataExpert-io/data-engineer-handbook** repository provides a complete Docker Compose stack for Apache Iceberg lakehouse experimentation with a single `make up` command.
- The architecture combines **Spark**, **Iceberg REST Catalog**, and **MinIO** to deliver ACID transactions and schema evolution capabilities using the `tabulario/spark-iceberg` and `tabulario/iceberg-rest` images.
- Configuration in [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) uses the **S3FileIO** implementation to persist data in the MinIO `warehouse` bucket with credentials shared across services.
- **Bucket partitioning** with `bucket(16, match_id)` optimizes join performance by distributing data across 16 buckets based on hash values.
- Access the demonstration notebook at `localhost:8888` to run the `bucket-joins-in-iceberg.ipynb` examples and validate the lakehouse setup.

## Frequently Asked Questions

### What is the role of the Iceberg REST Catalog in this architecture?

The Iceberg REST Catalog serves as the centralized metadata service that exposes table information via a REST API. According to the DataExpert-io/data-engineer-handbook source code, it runs the `tabulario/iceberg-rest` image and configures the `S3FileIO` implementation to persist metadata in the MinIO object store, enabling multiple Spark sessions to share table state consistently through the `rest` service defined in [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml).

### Why does the setup use MinIO instead of Amazon S3?

MinIO provides an S3-compatible object storage layer that runs locally without cloud credentials or costs. The [`docker-compose.yaml`](https://github.com/DataExpert-io/data-engineer-handbook/blob/main/docker-compose.yaml) configures MinIO with the `minio/minio` image and automatically creates the `warehouse` bucket via the `mc` sidecar container, making it ideal for development and testing Iceberg lakehouse patterns while maintaining full API compatibility with AWS S3.

### How does bucket partitioning improve query performance in Iceberg?

Bucket partitioning distributes data files across a fixed number of buckets based on a hash of the bucket column, in this case `bucket(16, match_id)`. When joining tables on the bucketed column, Spark can avoid shuffling data across the network by matching bucket indices directly, as demonstrated in the `bucket-joins-in-iceberg.ipynb` notebook where `spark.sql.autoBroadcastJoinThreshold` is set to `-1` to force this optimization and reveal the physical execution plan.

### Can I use PySpark instead of Scala with this Iceberg setup?

Yes, the Docker Compose environment supports both languages interchangeably. The PySpark configuration requires setting the `spark.sql.extensions` to `org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions` and configuring the `spark_catalog` to use the Hadoop catalog type pointing to `s3://warehouse/`, allowing you to execute identical `CREATE TABLE` and `SELECT` statements using Python syntax against the same Iceberg tables.