How to Set Up a Local OpenMetadata Instance for Testing Features

Run a complete local OpenMetadata stack with one command using the run_local_docker.sh helper script, which orchestrates Docker containers for the metadata store, search service, API server, and Airflow-driven ingestion pipeline.

OpenMetadata is a comprehensive metadata platform that unifies data discovery, lineage, and governance. For developers and data engineers who need to test features before production deployment, the project provides a fully containerized local environment. This guide walks through the exact steps implemented in docker/run_local_docker.sh and docker/run_local_docker_common.sh to launch all services with proper health checks and sample data.

Architecture of the Local OpenMetadata Stack

Understanding the component interactions helps troubleshoot setup issues and optimize resource allocation.

Component Role Container Name Default Configuration
MySQL or PostgreSQL Relational metadata store openmetadata_mysql or openmetadata_postgresql Defined in docker/development/docker-compose.yml or docker-compose-postgres.yml
Elasticsearch or OpenSearch Full-text search and lineage indexing openmetadata_elasticsearch or openmetadata_opensearch docker.elastic.co/elasticsearch/elasticsearch:9.3.0
OpenMetadata Server REST API, authentication, and UI openmetadata_server Built from docker/development/Dockerfile
Airflow Ingestion Orchestrates metadata extraction DAGs openmetadata_ingestion Built from ingestion/Dockerfile.ci

When startup completes, services are available at:

Prerequisites for Local OpenMetadata Setup

Before running the setup script, ensure your environment meets these requirements as specified in the repository's AGENTS.md:

  1. Docker Compose – Container orchestration engine
  2. Java 21 – For building the backend services
  3. Maven – Java build tool
  4. Python 3.10–3.12 – For ingestion framework
  5. Python virtual environment – Isolated dependencies for make install_dev generate

Step-by-Step: Running the Local OpenMetadata Setup Script

The docker/run_local_docker.sh script is a thin wrapper that sources docker/run_local_docker_common.sh and invokes run_local_docker_main(). The common script handles argument parsing (lines 32–44), orchestration, and health checks.

Basic Command Structure


# From repository root

./docker/run_local_docker.sh -m <mode> -d <database> -s <skip_maven> -i <ingestion> -r <reset_volumes>
Flag Values Description
-m ui (default) or no-ui Start with or without UI build
-d mysql (default) or postgresql Database backend
-s true or false (default) Skip Maven build (use existing artifacts)
-i true (default) or false Build ingestion container and generate Python models
-r true or false (default) Remove Docker volumes for fresh database
-x true or false (default) Expose JVM debug port 5005

Example: Complete Fresh Start with UI and Ingestion

./docker/run_local_docker.sh -m ui -d mysql -s false -i true -r true

This command:

  1. Builds the Java backend with Maven (-s false)
  2. Generates Python Pydantic models via make install_dev generate (lines 1–8 in run_local_docker_common.sh)
  3. Selects docker/development/docker-compose.yml for MySQL
  4. Removes existing volumes for clean state (-r true)
  5. Builds and starts all containers including Airflow ingestion

Example: Backend-Only Development Mode

./docker/run_local_docker.sh -m no-ui -d postgresql -s true -i false

This skips UI build, uses PostgreSQL, reuses existing Maven artifacts, and omits ingestion for faster startup.

Example: Enable JVM Debugging

./docker/run_local_docker.sh -x true

Exposes port 5005 for attaching a remote debugger to the OpenMetadata server.

What Happens During Script Execution

The run_local_docker_common.sh script performs these orchestration steps:

1. Argument Parsing and Validation

Lines 32–44 parse flags and set defaults. The run_local_docker_main() function validates combinations (e.g., no-ui mode with -s true requires pre-built JARs).

2. Maven Build (Unless Skipped)

When -s false, the script executes:

  • mvn clean package for full build
  • mvn clean package -pl !openmetadata-ui for no-ui mode

3. Python Model Generation

With -i true and active virtual environment:

make install_dev generate

This creates Pydantic models from the OpenMetadata JSON schemas.

4. Docker Compose Selection

Database Compose File
MySQL docker/development/docker-compose.yml
PostgreSQL docker/development/docker-compose-postgres.yml

5. Container Lifecycle

  • docker compose build – Builds custom images
  • docker compose up -d – Starts services detached
  • Volume removal with -r true for clean database state

6. Health Checks and Service Verification

The script polls services until ready:

  • Elasticsearch: curl http://localhost:9200/_cat/indices/...
  • OpenMetadata API: Authenticates and fetches tables via /api/v1/tables
  • Airflow (when -i true): Verifies /auth/token and DAG endpoints

7. Sample Data Ingestion

For testing features, the script:

  • Unpauses DAGs: sample_data, extended_sample_data, sample_usage, index_metadata, sample_lineage
  • Triggers sample_data DAG to populate the metadata store

8. Search Re-indexing

Finally, the script ensures SearchIndexingApplication is installed and triggers it, waiting for completion. This makes all ingested metadata searchable.

Manual Search Re-indexing

If you need to re-index without full restart:

./docker/run_local_docker.sh -m ui -d mysql -i false -r false \
    && ./docker/run_local_docker_common.sh trigger_app_and_wait "SearchIndexingApplication"

Troubleshooting Common Issues

Symptom Likely Cause Solution
Maven build fails Java version mismatch Verify Java 21 with java -version
make: command not found Missing build tools Install make and build-essential
Containers exit immediately Port conflicts Check ports 8585, 8080, 9200, 3306 are free
Health checks timeout Insufficient resources Allocate at least 8GB RAM to Docker
Ingestion DAGs fail Missing Python dependencies Ensure virtual environment is active before running script

Summary

Setting up a local OpenMetadata instance for testing requires these key steps:

  • Use docker/run_local_docker.sh with appropriate flags for your use case (UI mode, database choice, Maven skip, ingestion enable)
  • Meet prerequisites: Docker Compose, Java 21, Maven, Python 3.10–3.12 with virtual environment
  • The orchestration script handles building, container startup, health checks, sample data ingestion, and search re-indexing automatically
  • Access the UI at http://localhost:8585, API at http://localhost:8585/api, and Airflow at http://localhost:8080

Frequently Asked Questions

How long does the initial OpenMetadata local setup take?

First-time setup typically takes 15–30 minutes depending on your hardware and network. The Maven build of the Java backend is the longest step. Subsequent runs with -s true (skip Maven) reduce startup to 3–5 minutes.

Can I run OpenMetadata locally without the ingestion pipeline?

Yes. Pass -i false to disable ingestion container building and Python model generation. This accelerates startup when you only need to test UI features or API behavior without metadata extraction workflows.

What database should I choose for local testing?

MySQL is the default and most widely tested configuration. Use PostgreSQL with -d postgresql if your production environment requires it or if you encounter MySQL-specific issues. Both databases use officially supported Docker images with identical feature sets.

How do I preserve data between restarts?

Pass -r false to skip volume removal. This retains your MySQL or PostgreSQL data, Elasticsearch indices, and Airflow metadata across container restarts. Use this when testing incremental ingestion or maintaining consistent test datasets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →