How to Build OpenMetadata from Source: A Complete Developer Guide
Build OpenMetadata from source by setting up Java 21, Python 3.10/3.11, and Node/Yarn, then running make generate followed by ./docker/run_local_docker.sh to compile the full stack.
This guide walks you through building OpenMetadata from source code, covering the repository's three-layer build system: Python ingestion components, Java backend services, and React-based UI. Whether you're contributing features or deploying custom modifications, understanding how to build OpenMetadata from source ensures you can iterate efficiently on any part of the codebase.
Prerequisites for Building OpenMetadata from Source
Before compiling, verify your environment meets these requirements. The OpenMetadata repository requires specific versions across multiple technology stacks.
Required Tools and Versions
| Component | Version | Verification Command |
|---|---|---|
| Java | 21 | java -version |
| Maven | 3.8+ | mvn -v |
| Python | 3.10 or 3.11 | python3 --version |
| Node.js | 18+ | node -v |
| Yarn | 1.22+ | yarn -v |
| Docker | Latest | docker --version |
Python Virtual Environment Setup
Create and activate a repository-wide virtual environment. This environment is required for ingestion development and code generation.
python3.11 -m venv env
source env/bin/activate
On Windows, use env\Scripts\activate instead.
Installing Ingestion Dependencies
The ingestion module contains Python connectors, metadata models, and pipeline frameworks. Install it in editable mode with development dependencies.
Install Development Environment
cd ingestion
make install_dev_env
This Makefile target, defined in ingestion/Makefile, installs:
- Pydantic for data validation
- ANTLR runtime for parsing fully qualified names (FQN)
- Test frameworks (pytest, coverage)
- Code quality tools (black, isort, pylint)
Return to the repository root after installation:
cd ..
Code Generation: Models and Parsers
OpenMetadata uses code generation to maintain consistency between its JSON schema definitions and runtime code. The make generate command orchestrates this process.
Run the Generation Pipeline
make generate
This top-level Makefile target executes three operations in sequence:
- Pydantic model generation –
scripts/datamodel_generation.pyconverts JSON schemas inopenmetadata-spec/into Python classes - Python ANTLR parser – compiles FQN grammar for Python ingestion
- JavaScript ANTLR parser – compiles FQN grammar for UI consumption
Key Files in Code Generation
| File | Purpose |
|---|---|
Makefile |
Orchestrates all build targets |
scripts/datamodel_generation.py |
JSON schema → Pydantic conversion |
openmetadata-spec/ |
Source JSON schemas |
Building the Full OpenMetadata Stack
With dependencies installed and code generated, compile the complete application. The repository provides a unified script for this: docker/run_local_docker.sh.
First-Time Full Build
For initial builds or after Java/UI changes, run:
./docker/run_local_docker.sh \
-m ui \
-d mysql \
-s false \
-i true \
-r true
Flag Explanation
| Flag | Value | Meaning |
|---|---|---|
-m |
ui |
Build both Java backend and React UI via Maven |
-d |
mysql |
Use MySQL as metadata store (alternative: postgresql) |
-s |
false |
Do not skip Maven build (full compilation) |
-i |
true |
Build ingestion Docker image |
-r |
true |
Run sample data DAGs and re-index search after startup |
This script, located at docker/run_local_docker.sh, handles:
- Maven compilation of
openmetadata-service(Java backend) - Yarn build of
openmetadata-ui(React frontend) - Docker image construction for ingestion
- Docker Compose orchestration of all services
Fast Rebuild for Ingestion Development
When iterating on Python ingestion code only, skip the time-consuming Maven build.
Quick Iteration Command
./docker/run_local_docker.sh \
-m ui \
-d mysql \
-s true \
-i true \
-r false
The critical difference is -s true, which skips Maven compilation. This rebuilds only:
- Python ingestion Docker image (with your code changes)
- Docker Compose services
Typical iteration time drops from 10-15 minutes to 2-3 minutes.
Verifying Your Build
Once services start, confirm functionality through the web interface.
Access Points
| Service | URL | Purpose |
|---|---|---|
| OpenMetadata UI | http://localhost:8585 |
Main application |
| Airflow (optional) | http://localhost:8080 |
Pipeline orchestration |
Functional Verification
- Navigate to
http://localhost:8585 - Complete initial setup (admin user creation)
- Go to Settings → Services → Add a connector
- Test a simple connection (e.g., sample data)
Successful connector validation confirms the full stack is operational.
Stopping and Cleaning Up
When finished developing, cleanly shut down all services.
Graceful Shutdown
cd docker/development
docker compose down -v
The -v flag removes named volumes, ensuring a clean state for subsequent builds. Omit it to preserve data between sessions.
Summary
Building OpenMetadata from source requires coordinating three technology stacks:
- Python ingestion – Install via
make install_dev_enviningestion/, then runmake generatefor model generation - Java backend – Compiled via Maven through
docker/run_local_docker.shwith-s false - React UI – Bundled via Yarn through the same script with
-m ui
Key optimization: Use -s true for fast Python-only iterations, avoiding 10+ minute Maven builds. The docker/run_local_docker.sh script remains the central orchestration point for all full-stack operations.
Frequently Asked Questions
What Java version is required to build OpenMetadata from source?
OpenMetadata requires Java 21 for compilation. Verify your installation with java -version. The Maven build will fail with older versions, particularly versions below Java 17.
Can I use PostgreSQL instead of MySQL for the metadata store?
Yes. Change the -d flag in run_local_docker.sh from mysql to postgresql. The script automatically configures the appropriate Docker Compose services and connection strings for your chosen database.
How long does a full build take, and how can I speed it up?
A full build with Maven compilation takes 10-15 minutes depending on hardware. For rapid iteration on Python ingestion code only, use -s true to skip Maven, reducing build time to 2-3 minutes. The ingestion Docker image rebuilds with your Python changes while reusing existing Java and UI artifacts.
Where are the generated Pydantic models created, and why are they needed?
The make generate command creates Pydantic models from JSON schemas in openmetadata-spec/ using scripts/datamodel_generation.py. These models provide type-safe Python representations of OpenMetadata's entity definitions, enabling IDE autocomplete and runtime validation for ingestion connectors and API clients.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →