How to Build OpenMetadata from Source: A Complete Developer Guide

Build OpenMetadata from source by setting up Java 21, Python 3.10/3.11, and Node/Yarn, then running make generate followed by ./docker/run_local_docker.sh to compile the full stack.

This guide walks you through building OpenMetadata from source code, covering the repository's three-layer build system: Python ingestion components, Java backend services, and React-based UI. Whether you're contributing features or deploying custom modifications, understanding how to build OpenMetadata from source ensures you can iterate efficiently on any part of the codebase.


Prerequisites for Building OpenMetadata from Source

Before compiling, verify your environment meets these requirements. The OpenMetadata repository requires specific versions across multiple technology stacks.

Required Tools and Versions

Component Version Verification Command
Java 21 java -version
Maven 3.8+ mvn -v
Python 3.10 or 3.11 python3 --version
Node.js 18+ node -v
Yarn 1.22+ yarn -v
Docker Latest docker --version

Python Virtual Environment Setup

Create and activate a repository-wide virtual environment. This environment is required for ingestion development and code generation.

python3.11 -m venv env
source env/bin/activate

On Windows, use env\Scripts\activate instead.


Installing Ingestion Dependencies

The ingestion module contains Python connectors, metadata models, and pipeline frameworks. Install it in editable mode with development dependencies.

Install Development Environment

cd ingestion
make install_dev_env

This Makefile target, defined in ingestion/Makefile, installs:

  • Pydantic for data validation
  • ANTLR runtime for parsing fully qualified names (FQN)
  • Test frameworks (pytest, coverage)
  • Code quality tools (black, isort, pylint)

Return to the repository root after installation:

cd ..

Code Generation: Models and Parsers

OpenMetadata uses code generation to maintain consistency between its JSON schema definitions and runtime code. The make generate command orchestrates this process.

Run the Generation Pipeline

make generate

This top-level Makefile target executes three operations in sequence:

  1. Pydantic model generation – scripts/datamodel_generation.py converts JSON schemas in openmetadata-spec/ into Python classes
  2. Python ANTLR parser – compiles FQN grammar for Python ingestion
  3. JavaScript ANTLR parser – compiles FQN grammar for UI consumption

Key Files in Code Generation

File Purpose
Makefile Orchestrates all build targets
scripts/datamodel_generation.py JSON schema → Pydantic conversion
openmetadata-spec/ Source JSON schemas

Building the Full OpenMetadata Stack

With dependencies installed and code generated, compile the complete application. The repository provides a unified script for this: docker/run_local_docker.sh.

First-Time Full Build

For initial builds or after Java/UI changes, run:

./docker/run_local_docker.sh \
  -m ui \
  -d mysql \
  -s false \
  -i true \
  -r true

Flag Explanation

Flag Value Meaning
-m ui Build both Java backend and React UI via Maven
-d mysql Use MySQL as metadata store (alternative: postgresql)
-s false Do not skip Maven build (full compilation)
-i true Build ingestion Docker image
-r true Run sample data DAGs and re-index search after startup

This script, located at docker/run_local_docker.sh, handles:

  • Maven compilation of openmetadata-service (Java backend)
  • Yarn build of openmetadata-ui (React frontend)
  • Docker image construction for ingestion
  • Docker Compose orchestration of all services

Fast Rebuild for Ingestion Development

When iterating on Python ingestion code only, skip the time-consuming Maven build.

Quick Iteration Command

./docker/run_local_docker.sh \
  -m ui \
  -d mysql \
  -s true \
  -i true \
  -r false

The critical difference is -s true, which skips Maven compilation. This rebuilds only:

  • Python ingestion Docker image (with your code changes)
  • Docker Compose services

Typical iteration time drops from 10-15 minutes to 2-3 minutes.


Verifying Your Build

Once services start, confirm functionality through the web interface.

Access Points

Service URL Purpose
OpenMetadata UI http://localhost:8585 Main application
Airflow (optional) http://localhost:8080 Pipeline orchestration

Functional Verification

  1. Navigate to http://localhost:8585
  2. Complete initial setup (admin user creation)
  3. Go to Settings → Services → Add a connector
  4. Test a simple connection (e.g., sample data)

Successful connector validation confirms the full stack is operational.


Stopping and Cleaning Up

When finished developing, cleanly shut down all services.

Graceful Shutdown

cd docker/development
docker compose down -v

The -v flag removes named volumes, ensuring a clean state for subsequent builds. Omit it to preserve data between sessions.


Summary

Building OpenMetadata from source requires coordinating three technology stacks:

  • Python ingestion – Install via make install_dev_env in ingestion/, then run make generate for model generation
  • Java backend – Compiled via Maven through docker/run_local_docker.sh with -s false
  • React UI – Bundled via Yarn through the same script with -m ui

Key optimization: Use -s true for fast Python-only iterations, avoiding 10+ minute Maven builds. The docker/run_local_docker.sh script remains the central orchestration point for all full-stack operations.


Frequently Asked Questions

What Java version is required to build OpenMetadata from source?

OpenMetadata requires Java 21 for compilation. Verify your installation with java -version. The Maven build will fail with older versions, particularly versions below Java 17.

Can I use PostgreSQL instead of MySQL for the metadata store?

Yes. Change the -d flag in run_local_docker.sh from mysql to postgresql. The script automatically configures the appropriate Docker Compose services and connection strings for your chosen database.

How long does a full build take, and how can I speed it up?

A full build with Maven compilation takes 10-15 minutes depending on hardware. For rapid iteration on Python ingestion code only, use -s true to skip Maven, reducing build time to 2-3 minutes. The ingestion Docker image rebuilds with your Python changes while reusing existing Java and UI artifacts.

Where are the generated Pydantic models created, and why are they needed?

The make generate command creates Pydantic models from JSON schemas in openmetadata-spec/ using scripts/datamodel_generation.py. These models provide type-safe Python representations of OpenMetadata's entity definitions, enabling IDE autocomplete and runtime validation for ingestion connectors and API clients.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →