# How to Contribute to the OpenMetadata Project: A Complete Developer Guide

> Learn how to contribute to OpenMetadata. Set up your environment, follow design patterns, write tests, and submit pull requests to join our open-source data cataloging project and shape its future.

- Repository: [OpenMetadata/OpenMetadata](https://github.com/open-metadata/OpenMetadata)
- Tags: how-to-guide
- Published: 2026-04-23

---

**Contributing to OpenMetadata requires setting up a multi-language development environment (Java 21, Node 18+, Python 3.10-3.12), following schema-first design patterns, writing tests before code, and submitting pull requests through the standard GitHub workflow.**

OpenMetadata is a unified metadata platform for data discovery, observability, and governance. Learning how to contribute to this open-source project opens doors to collaborating on a modern data stack used by organizations worldwide. The repository follows a **schema-first, test-driven** workflow across four main modules: Java backend services, React/TypeScript UI, Python ingestion framework, and shared JSON schemas.

---

## Setting Up Your Local Development Environment

Before writing any code, you must configure the correct toolchain for your target module. OpenMetadata's multi-language architecture requires different dependencies depending on where you plan to contribute.

### Java Backend Setup

The backend services require **JDK 21** and **Maven**. Initialize your environment with:

```bash

# Install JDK 21 and Maven first, then:

make install_dev_env

```

This command pulls Java dependencies and configures the build system. The backend source lives in `openmetadata-service/src/main/java/` and follows specific architectural patterns documented in [`DEVELOPER.md`](https://github.com/open-metadata/OpenMetadata/blob/main/DEVELOPER.md).

### React/TypeScript UI Setup

The frontend requires **Node.js 18 or higher** and **Yarn**. Navigate to the UI directory and install dependencies:

```bash
cd openmetadata-ui/src/main/resources/ui
yarn

```

The UI is built from `openmetadata-ui-core-components` and auto-generates TypeScript types from JSON schemas. Always import types from `openmetadata-ui/src/main/resources/ui/src/generated/` rather than defining them manually.

### Python Ingestion Framework Setup

The ingestion connectors require **Python 3.10, 3.11, or 3.12**. Create a virtual environment and run:

```bash

# In the ingestion/ folder

make install_dev_env

```

This installs the `metadata` package and all connector dependencies. The ingestion framework uses a **topology-based extraction pattern** defined in `ingestion/src/metadata/ingestion/source/`.

### Schema-Driven Code Generation

OpenMetadata uses a **schema-first design** where all entities are defined in `openmetadata-spec/src/main/resources/json/schema/`. After editing any JSON schema, regenerate dependent code:

```bash
make generate

```

This single command updates Java POJOs, Python models, and TypeScript types across all modules—maintaining type safety and consistency.

---

## Finding Your First Contribution

The OpenMetadata project uses GitHub issues to track work. Filter by the **"good-first-issue"** label to find beginner-friendly tasks ranging from typo fixes to small feature additions.

For more substantial contributions:

- **New connectors** – Add service-specific connectors under `ingestion/src/metadata/ingestion/source/`. The [`DEVELOPER.md`](https://github.com/open-metadata/OpenMetadata/blob/main/DEVELOPER.md) file (lines 71-82) contains a detailed checklist.
- **Bug fixes** – Search the codebase using `grep` or GitHub's search to identify the relevant module (backend, UI, or Python).
- **Feature development** – Discuss significant changes in a GitHub issue before investing implementation time.

---

## Understanding OpenMetadata's Architecture

Contributing effectively requires following established patterns in each module. The codebase enforces consistency through base classes and code generation.

### Backend Architecture Patterns

The Java backend follows a layered architecture with clear separation of concerns:

**Entity Registry**

All entity constants are centralized in [`openmetadata-service/src/main/java/org/openmetadata/schema/type/Entity.java`](https://github.com/open-metadata/OpenMetadata/blob/main/openmetadata-service/src/main/java/org/openmetadata/schema/type/Entity.java). When adding a new entity, you must register it here:

```java
// From Entity.java
Entity.TABLE, Entity.DASHBOARD, Entity.MY_ENTITY  // Your new entity

```

**Repository Layer**

Data access extends `EntityRepository<E>` with JDBI3. Required methods include `setFullyQualifiedName`, `prepare`, `storeEntity`, and `storeRelationships`. The base class provides pagination, field filtering, and relationship management.

**REST Resource Layer**

API endpoints inherit from `EntityResource<E, R extends EntityRepository<E>>`. This provides automatic CRUD operations, pagination, and field filtering without boilerplate code.

### Frontend Type Safety

The UI enforces strict TypeScript discipline through generated types:

- Import all entity types from `openmetadata-ui/src/main/resources/ui/src/generated/`
- Use components from `openmetadata-ui-core-components` for consistency
- Localize strings through `useTranslation()` with keys in [`locale/languages/en-us.json`](https://github.com/open-metadata/OpenMetadata/blob/main/locale/languages/en-us.json)

### Python Ingestion Topology

Connectors implement a **ServiceTopology** composed of **TopologyNode** elements. This declarative pattern drives depth-first extraction and handles complex dependency graphs between metadata objects automatically.

---

## Writing Tests Before Code

OpenMetadata enforces **test-driven development (TDD)** for every change. The CI pipeline blocks merges without adequate test coverage.

| Language | Test Framework | Command |
|----------|--------------|---------|
| Java | JUnit + integration tests | `mvn verify` |
| TypeScript | Jest (unit) + Playwright (E2E) | `yarn test` and `yarn playwright:run` |
| Python | pytest | `make unit_ingestion` or `pytest` |

For backend entities, create a `*IT.java` integration test that exercises the full REST flow:

```java
public class MyEntityIT extends BaseEntityIT<MyEntity, CreateMyEntity> {
  @Override
  public CreateMyEntity createRequest(String name) {
    return new CreateMyEntity().withName(name).withDescription("Demo");
  }

  @Override
  public void validateCreatedEntity(MyEntity entity, CreateMyEntity request) {
    assertEquals(request.getName(), entity.getName());
    assertEquals(request.getDescription(), entity.getDescription());
  }
}

```

---

## Running Quality Checks and Submitting Your Contribution

Before opening a pull request, ensure all automated checks pass:

### Java Quality Commands

```bash
mvn spotless:apply   # Auto-format code

mvn verify           # Run unit and integration tests

```

### TypeScript Quality Commands

```bash
yarn lint:fix        # Fix linting issues

yarn test            # Jest unit tests

yarn playwright:run  # End-to-end tests

```

### Python Quality Commands

```bash
make py_format       # Auto-format Python code

make lint            # Run pylint and mypy

pytest               # Unit tests

```

All checks are enforced in CI—pull requests cannot merge without passing the full pipeline.

### Pull Request Workflow

1. **Fork** the repository and create a feature branch: `git checkout -b my-feature`
2. **Commit** using the message convention described in [`CONTRIBUTING.md`](https://github.com/open-metadata/OpenMetadata/blob/main/CONTRIBUTING.md)
3. **Push** your branch and open a **GitHub pull request**, linking any related issues
4. **Respond** to reviewer feedback—the team expects tests, lint passes, and documentation updates for every change

---

## Code Example: Adding a Backend Entity

Here's a complete example implementing a new entity in the OpenMetadata backend, following the patterns established in `EntityResource` and `EntityRepository`:

```java
// src/main/java/org/openmetadata/service/resources/myentity/MyEntityResource.java
package org.openmetadata.service.resources.myentity;

import org.openmetadata.schema.entity.services.MyEntity;
import org.openmetadata.service.resources.Collection;
import org.openmetadata.service.resources.EntityResource;
import org.openmetadata.service.security.Authorizer;
import org.openmetadata.service.util.Limits;

import javax.ws.rs.*;
import javax.ws.rs.core.MediaType;

import org.eclipse.microprofile.openapi.annotations.tags.Tag;

@Path("/v1/myEntities")
@Tag(name = "MyEntities")
@Consumes(MediaType.APPLICATION_JSON)
@Produces(MediaType.APPLICATION_JSON)
@Collection(name = "myEntities")
public class MyEntityResource extends EntityResource<MyEntity, MyEntityRepository> {
  public static final String COLLECTION_PATH = "/v1/myEntities/";
  public static final String FIELDS = "owners,tags,domain";

  public MyEntityResource(Authorizer authorizer, Limits limits) {
    super(Entity.MY_ENTITY, authorizer, limits);
  }

  // CRUD endpoints inherited from EntityResource—no additional code required
  public static class MyEntityList extends ResultList<MyEntity> {}
}

```

**Key implementation points:**

- Extends `EntityResource` to automatically expose CRUD, pagination, and field filtering
- Declares `COLLECTION_PATH` and `FIELDS` constants as required by the base class
- Uses `Entity.MY_ENTITY` constant registered in [`Entity.java`](https://github.com/open-metadata/OpenMetadata/blob/main/Entity.java)
- Annotates with `@Collection` for automatic endpoint registration

---

## Summary

- **Set up** the correct development environment for your target module: Java 21 for backend, Node 18+ for UI, Python 3.10-3.12 for ingestion
- **Follow** schema-first design by editing JSON schemas in `openmetadata-spec/` and running `make generate`
- **Use** established architectural patterns: `EntityResource` and `EntityRepository` for backend, generated types for UI, `ServiceTopology` for Python connectors
- **Write** tests before implementation using JUnit, Jest, or pytest depending on your module
- **Pass** all quality checks with `mvn verify`, `yarn test`, or `make lint` before submitting
- **Submit** pull requests through the standard GitHub fork-and-branch workflow with proper commit conventions

---

## Frequently Asked Questions

### What programming languages does OpenMetadata use?

OpenMetadata is a polyglot codebase using **Java 21** for the backend services, **TypeScript/React** for the UI, and **Python 3.10-3.12** for the metadata ingestion framework. Contributors typically specialize in one module but should understand how JSON schemas tie all three together.

### Do I need to write tests for my contribution?

**Yes—test-driven development is mandatory.** The CI pipeline enforces this through `/test-enforcement` policies. Java contributions need JUnit integration tests (`*IT.java`), TypeScript requires Jest unit tests and Playwright E2E tests, and Python uses pytest. Pull requests fail without adequate coverage.

### How does the schema-first design work?

All entities are defined as **JSON schemas** in `openmetadata-spec/src/main/resources/json/schema/`. When you modify a schema, running `make generate` automatically produces Java POJOs, Python models, and TypeScript types. This ensures type consistency across all modules without manual synchronization.

### What makes a good first contribution?

Start with issues labeled **"good-first-issue"** on the GitHub tracker. These range from documentation fixes to small bug fixes. Alternatively, adding a new connector follows a well-documented checklist in [`DEVELOPER.md`](https://github.com/open-metadata/OpenMetadata/blob/main/DEVELOPER.md) (lines 71-82) and provides structured guidance for first-time contributors.