How to Contribute to the OpenMetadata Project: A Complete Developer Guide

Contributing to OpenMetadata requires setting up a multi-language development environment (Java 21, Node 18+, Python 3.10-3.12), following schema-first design patterns, writing tests before code, and submitting pull requests through the standard GitHub workflow.

OpenMetadata is a unified metadata platform for data discovery, observability, and governance. Learning how to contribute to this open-source project opens doors to collaborating on a modern data stack used by organizations worldwide. The repository follows a schema-first, test-driven workflow across four main modules: Java backend services, React/TypeScript UI, Python ingestion framework, and shared JSON schemas.


Setting Up Your Local Development Environment

Before writing any code, you must configure the correct toolchain for your target module. OpenMetadata's multi-language architecture requires different dependencies depending on where you plan to contribute.

Java Backend Setup

The backend services require JDK 21 and Maven. Initialize your environment with:


# Install JDK 21 and Maven first, then:

make install_dev_env

This command pulls Java dependencies and configures the build system. The backend source lives in openmetadata-service/src/main/java/ and follows specific architectural patterns documented in DEVELOPER.md.

React/TypeScript UI Setup

The frontend requires Node.js 18 or higher and Yarn. Navigate to the UI directory and install dependencies:

cd openmetadata-ui/src/main/resources/ui
yarn

The UI is built from openmetadata-ui-core-components and auto-generates TypeScript types from JSON schemas. Always import types from openmetadata-ui/src/main/resources/ui/src/generated/ rather than defining them manually.

Python Ingestion Framework Setup

The ingestion connectors require Python 3.10, 3.11, or 3.12. Create a virtual environment and run:


# In the ingestion/ folder

make install_dev_env

This installs the metadata package and all connector dependencies. The ingestion framework uses a topology-based extraction pattern defined in ingestion/src/metadata/ingestion/source/.

Schema-Driven Code Generation

OpenMetadata uses a schema-first design where all entities are defined in openmetadata-spec/src/main/resources/json/schema/. After editing any JSON schema, regenerate dependent code:

make generate

This single command updates Java POJOs, Python models, and TypeScript types across all modules—maintaining type safety and consistency.


Finding Your First Contribution

The OpenMetadata project uses GitHub issues to track work. Filter by the "good-first-issue" label to find beginner-friendly tasks ranging from typo fixes to small feature additions.

For more substantial contributions:

  • New connectors – Add service-specific connectors under ingestion/src/metadata/ingestion/source/. The DEVELOPER.md file (lines 71-82) contains a detailed checklist.
  • Bug fixes – Search the codebase using grep or GitHub's search to identify the relevant module (backend, UI, or Python).
  • Feature development – Discuss significant changes in a GitHub issue before investing implementation time.

Understanding OpenMetadata's Architecture

Contributing effectively requires following established patterns in each module. The codebase enforces consistency through base classes and code generation.

Backend Architecture Patterns

The Java backend follows a layered architecture with clear separation of concerns:

Entity Registry

All entity constants are centralized in openmetadata-service/src/main/java/org/openmetadata/schema/type/Entity.java. When adding a new entity, you must register it here:

// From Entity.java
Entity.TABLE, Entity.DASHBOARD, Entity.MY_ENTITY  // Your new entity

Repository Layer

Data access extends EntityRepository<E> with JDBI3. Required methods include setFullyQualifiedName, prepare, storeEntity, and storeRelationships. The base class provides pagination, field filtering, and relationship management.

REST Resource Layer

API endpoints inherit from EntityResource<E, R extends EntityRepository<E>>. This provides automatic CRUD operations, pagination, and field filtering without boilerplate code.

Frontend Type Safety

The UI enforces strict TypeScript discipline through generated types:

  • Import all entity types from openmetadata-ui/src/main/resources/ui/src/generated/
  • Use components from openmetadata-ui-core-components for consistency
  • Localize strings through useTranslation() with keys in locale/languages/en-us.json

Python Ingestion Topology

Connectors implement a ServiceTopology composed of TopologyNode elements. This declarative pattern drives depth-first extraction and handles complex dependency graphs between metadata objects automatically.


Writing Tests Before Code

OpenMetadata enforces test-driven development (TDD) for every change. The CI pipeline blocks merges without adequate test coverage.

Language Test Framework Command
Java JUnit + integration tests mvn verify
TypeScript Jest (unit) + Playwright (E2E) yarn test and yarn playwright:run
Python pytest make unit_ingestion or pytest

For backend entities, create a *IT.java integration test that exercises the full REST flow:

public class MyEntityIT extends BaseEntityIT<MyEntity, CreateMyEntity> {
  @Override
  public CreateMyEntity createRequest(String name) {
    return new CreateMyEntity().withName(name).withDescription("Demo");
  }

  @Override
  public void validateCreatedEntity(MyEntity entity, CreateMyEntity request) {
    assertEquals(request.getName(), entity.getName());
    assertEquals(request.getDescription(), entity.getDescription());
  }
}

Running Quality Checks and Submitting Your Contribution

Before opening a pull request, ensure all automated checks pass:

Java Quality Commands

mvn spotless:apply   # Auto-format code

mvn verify           # Run unit and integration tests

TypeScript Quality Commands

yarn lint:fix        # Fix linting issues

yarn test            # Jest unit tests

yarn playwright:run  # End-to-end tests

Python Quality Commands

make py_format       # Auto-format Python code

make lint            # Run pylint and mypy

pytest               # Unit tests

All checks are enforced in CI—pull requests cannot merge without passing the full pipeline.

Pull Request Workflow

  1. Fork the repository and create a feature branch: git checkout -b my-feature
  2. Commit using the message convention described in CONTRIBUTING.md
  3. Push your branch and open a GitHub pull request, linking any related issues
  4. Respond to reviewer feedback—the team expects tests, lint passes, and documentation updates for every change

Code Example: Adding a Backend Entity

Here's a complete example implementing a new entity in the OpenMetadata backend, following the patterns established in EntityResource and EntityRepository:

// src/main/java/org/openmetadata/service/resources/myentity/MyEntityResource.java
package org.openmetadata.service.resources.myentity;

import org.openmetadata.schema.entity.services.MyEntity;
import org.openmetadata.service.resources.Collection;
import org.openmetadata.service.resources.EntityResource;
import org.openmetadata.service.security.Authorizer;
import org.openmetadata.service.util.Limits;

import javax.ws.rs.*;
import javax.ws.rs.core.MediaType;

import org.eclipse.microprofile.openapi.annotations.tags.Tag;

@Path("/v1/myEntities")
@Tag(name = "MyEntities")
@Consumes(MediaType.APPLICATION_JSON)
@Produces(MediaType.APPLICATION_JSON)
@Collection(name = "myEntities")
public class MyEntityResource extends EntityResource<MyEntity, MyEntityRepository> {
  public static final String COLLECTION_PATH = "/v1/myEntities/";
  public static final String FIELDS = "owners,tags,domain";

  public MyEntityResource(Authorizer authorizer, Limits limits) {
    super(Entity.MY_ENTITY, authorizer, limits);
  }

  // CRUD endpoints inherited from EntityResource—no additional code required
  public static class MyEntityList extends ResultList<MyEntity> {}
}

Key implementation points:

  • Extends EntityResource to automatically expose CRUD, pagination, and field filtering
  • Declares COLLECTION_PATH and FIELDS constants as required by the base class
  • Uses Entity.MY_ENTITY constant registered in Entity.java
  • Annotates with @Collection for automatic endpoint registration

Summary

  • Set up the correct development environment for your target module: Java 21 for backend, Node 18+ for UI, Python 3.10-3.12 for ingestion
  • Follow schema-first design by editing JSON schemas in openmetadata-spec/ and running make generate
  • Use established architectural patterns: EntityResource and EntityRepository for backend, generated types for UI, ServiceTopology for Python connectors
  • Write tests before implementation using JUnit, Jest, or pytest depending on your module
  • Pass all quality checks with mvn verify, yarn test, or make lint before submitting
  • Submit pull requests through the standard GitHub fork-and-branch workflow with proper commit conventions

Frequently Asked Questions

What programming languages does OpenMetadata use?

OpenMetadata is a polyglot codebase using Java 21 for the backend services, TypeScript/React for the UI, and Python 3.10-3.12 for the metadata ingestion framework. Contributors typically specialize in one module but should understand how JSON schemas tie all three together.

Do I need to write tests for my contribution?

Yes—test-driven development is mandatory. The CI pipeline enforces this through /test-enforcement policies. Java contributions need JUnit integration tests (*IT.java), TypeScript requires Jest unit tests and Playwright E2E tests, and Python uses pytest. Pull requests fail without adequate coverage.

How does the schema-first design work?

All entities are defined as JSON schemas in openmetadata-spec/src/main/resources/json/schema/. When you modify a schema, running make generate automatically produces Java POJOs, Python models, and TypeScript types. This ensures type consistency across all modules without manual synchronization.

What makes a good first contribution?

Start with issues labeled "good-first-issue" on the GitHub tracker. These range from documentation fixes to small bug fixes. Alternatively, adding a new connector follows a well-documented checklist in DEVELOPER.md (lines 71-82) and provides structured guidance for first-time contributors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →