How to Set Up OpenMetadata Locally for Development: Complete Guide
Use the run_local_docker.sh script to spin up the full OpenMetadata development stack—including Java backend, React UI, Python ingestion, and all Docker services—with a single command.
OpenMetadata is a comprehensive metadata platform with four core components: metadata schemas, a Java-based backend service, a React frontend, and a Python ingestion framework. Setting up OpenMetadata locally for development requires coordinating these multi-language modules alongside Docker services for MySQL/PostgreSQL, Elasticsearch/OpenSearch, and Airflow. This guide walks through the complete setup using the official automation scripts and manual configuration options.
Prerequisites for OpenMetadata Local Development
Before running OpenMetadata locally, ensure your system meets these minimum requirements:
| Tool | Minimum Version | Purpose |
|---|---|---|
| Git | any | Clone the repository |
| Docker Compose | 2.20+ | Orchestrate database, search, and Airflow containers |
| Java | 21 | Build and run the Dropwizard backend |
| Node.js + Yarn | 18+ | Compile and serve the React UI |
| Python | 3.10 or 3.11 | Run ingestion framework and generate Pydantic models |
| Maven | 3.9+ | Backend build automation |
| Make | any | Convenience targets like make install_dev |
The repository includes a CLAUDE.md file that provides architectural context for these components. Understanding how metadata schemas, the backend store, APIs, and ingestion interact helps determine which modules need rebuilding during iterative development.
Step 1: Clone the OpenMetadata Repository
Start by cloning the official repository and navigating to the project root:
git clone https://github.com/open-metadata/OpenMetadata.git
cd OpenMetadata
This creates the complete project structure including openmetadata-service/ for the Java backend, openmetadata-ui/ for the React frontend, ingestion/ for Python connectors, and docker/ for orchestration scripts.
Step 2: Set Up the Python Ingestion Environment
The ingestion framework requires a dedicated virtual environment and Pydantic model generation from JSON schemas:
python3.11 -m venv env
source env/bin/activate
cd ingestion
make install_dev_env
make generate
cd ..
The make generate command is essential—it processes the canonical JSON schemas in openmetadata-spec/ to produce Python Pydantic models. Without this step, ingestion connectors cannot validate metadata objects against the platform's type system.
Step 3: Build the Java Backend (Optional)
For backend development, use Maven to compile the Dropwizard service. Two build modes are available:
# Fast build—skip tests and exclude UI module
mvn -DskipTests -DonlyBackend clean package -pl !openmetadata-ui
# Full build—include all modules
mvn -DskipTests clean package
The -DonlyBackend flag accelerates iteration when UI changes aren't needed. The run_local_docker.sh script automatically selects the appropriate Maven profile based on its -m flag.
Step 4: Launch the Complete Development Stack
The run_local_docker.sh script automates the entire orchestration process:
# Default: UI + MySQL + Elasticsearch + full ingestion
./docker/run_local_docker.sh
# Use PostgreSQL instead of MySQL
./docker/run_local_docker.sh -m ui -d postgresql
# Skip Maven build (use existing compiled artifacts)
./docker/run_local_docker.sh -s true
# Exclude ingestion services (faster startup)
./docker/run_local_docker.sh -i false
The script executes these operations in docker/run_local_docker_common.sh:
- Stops any existing containers and cleans up state
- Optionally runs Maven compilation
- Generates Pydantic models if ingestion is enabled
- Builds Docker images with proper service dependencies
- Starts containers for database, search, Airflow, and the OpenMetadata server
- Polls Elasticsearch/OpenSearch for cluster readiness
- Obtains Airflow authentication token and triggers the
sample_dataDAG - Validates successful data ingestion with timeout handling
- Installs and executes the
SearchIndexingApplication
Successful completion displays: ✔ OpenMetadata is up and running
Step 5: Access the Development Interfaces
Once running, access these local endpoints:
- OpenMetadata UI: http://localhost:8585
- Airflow UI: http://localhost:8080 (credentials:
admin/admin)
The React UI automatically receives JWT tokens for API authentication through the Docker entrypoint configuration. No manual login configuration is required for local development.
Development Workflow Optimization
Use these targeted commands for efficient iteration:
| Task | Command |
|---|---|
| Hot-reload UI changes | cd openmetadata-ui/src/main/resources/ui && yarn start |
| Restart backend after Java changes | ./docker/run_local_docker.sh -s true |
| Regenerate models after schema changes | cd ingestion && make generate && ./docker/run_local_docker.sh -i true |
| Run only Airflow without full stack | ./docker/run_local_docker.sh -i false then trigger DAGs manually |
| Execute unit tests | mvn test (Java), yarn test (UI), make unit_ingestion (Python) |
Cleanup and Reset
To completely remove all containers and persisted data:
./docker/run_local_docker.sh -r true
The -r flag deletes the Docker volume directory, ensuring no stale database state persists between development sessions.
Summary
- Prerequisites: Install Java 21, Python 3.10/3.11, Node 18+, Docker Compose 2.20+, Maven 3.9+, and Make
- Quick start: Run
./docker/run_local_docker.shafter cloning and setting up the Python virtual environment - Key script:
run_local_docker.shorchestrates Maven builds, Docker Compose, Pydantic model generation, and Airflow DAG execution - Access points: UI at
localhost:8585, Airflow atlocalhost:8080 - Optimization: Use
-s trueto skip Maven builds,-i falseto exclude ingestion, andyarn startfor UI hot-reload
Frequently Asked Questions
What is the fastest way to set up OpenMetadata locally for development?
The fastest path is running ./docker/run_local_docker.sh after completing the one-time Python environment setup. This single command builds all components, starts Docker services, generates Pydantic models, and seeds sample data. For subsequent iterations, add -s true to skip Maven compilation and -i false to exclude ingestion services.
Can I use PostgreSQL instead of MySQL for local development?
Yes. Pass the -d postgresql flag to run_local_docker.sh along with -m ui to specify the build mode. The script automatically selects docker/development/docker-compose-postgres.yml instead of the default MySQL-based compose file. All other setup steps remain identical.
Why do I need to run make generate before starting the services?
The make generate command creates Python Pydantic models from the canonical JSON schemas in openmetadata-spec/. These models define the metadata types used throughout the ingestion framework. Without generated models, Python connectors cannot validate or serialize metadata objects, causing runtime failures. The run_local_docker.sh script runs this automatically when ingestion is enabled.
How do I clean up and reset my local development environment?
Run ./docker/run_local_docker.sh -r true to destroy all containers, delete persisted volumes, and remove the Docker volume directory. This ensures no stale database state or cached artifacts remain. Alternatively, run docker compose down -v manually from the docker/development/ directory, then delete the docker-volume/ folder if it exists.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →