How to Contribute to the OpenMetadata Project: A Complete Developer Guide
Contributing to OpenMetadata requires setting up a multi-language development environment (Java 21, Node 18+, Python 3.10-3.12), following schema-first design patterns, writing tests before code, and submitting pull requests through the standard GitHub workflow.
OpenMetadata is a unified metadata platform for data discovery, observability, and governance. Learning how to contribute to this open-source project opens doors to collaborating on a modern data stack used by organizations worldwide. The repository follows a schema-first, test-driven workflow across four main modules: Java backend services, React/TypeScript UI, Python ingestion framework, and shared JSON schemas.
Setting Up Your Local Development Environment
Before writing any code, you must configure the correct toolchain for your target module. OpenMetadata's multi-language architecture requires different dependencies depending on where you plan to contribute.
Java Backend Setup
The backend services require JDK 21 and Maven. Initialize your environment with:
# Install JDK 21 and Maven first, then:
make install_dev_env
This command pulls Java dependencies and configures the build system. The backend source lives in openmetadata-service/src/main/java/ and follows specific architectural patterns documented in DEVELOPER.md.
React/TypeScript UI Setup
The frontend requires Node.js 18 or higher and Yarn. Navigate to the UI directory and install dependencies:
cd openmetadata-ui/src/main/resources/ui
yarn
The UI is built from openmetadata-ui-core-components and auto-generates TypeScript types from JSON schemas. Always import types from openmetadata-ui/src/main/resources/ui/src/generated/ rather than defining them manually.
Python Ingestion Framework Setup
The ingestion connectors require Python 3.10, 3.11, or 3.12. Create a virtual environment and run:
# In the ingestion/ folder
make install_dev_env
This installs the metadata package and all connector dependencies. The ingestion framework uses a topology-based extraction pattern defined in ingestion/src/metadata/ingestion/source/.
Schema-Driven Code Generation
OpenMetadata uses a schema-first design where all entities are defined in openmetadata-spec/src/main/resources/json/schema/. After editing any JSON schema, regenerate dependent code:
make generate
This single command updates Java POJOs, Python models, and TypeScript types across all modules—maintaining type safety and consistency.
Finding Your First Contribution
The OpenMetadata project uses GitHub issues to track work. Filter by the "good-first-issue" label to find beginner-friendly tasks ranging from typo fixes to small feature additions.
For more substantial contributions:
- New connectors – Add service-specific connectors under
ingestion/src/metadata/ingestion/source/. TheDEVELOPER.mdfile (lines 71-82) contains a detailed checklist. - Bug fixes – Search the codebase using
grepor GitHub's search to identify the relevant module (backend, UI, or Python). - Feature development – Discuss significant changes in a GitHub issue before investing implementation time.
Understanding OpenMetadata's Architecture
Contributing effectively requires following established patterns in each module. The codebase enforces consistency through base classes and code generation.
Backend Architecture Patterns
The Java backend follows a layered architecture with clear separation of concerns:
Entity Registry
All entity constants are centralized in openmetadata-service/src/main/java/org/openmetadata/schema/type/Entity.java. When adding a new entity, you must register it here:
// From Entity.java
Entity.TABLE, Entity.DASHBOARD, Entity.MY_ENTITY // Your new entity
Repository Layer
Data access extends EntityRepository<E> with JDBI3. Required methods include setFullyQualifiedName, prepare, storeEntity, and storeRelationships. The base class provides pagination, field filtering, and relationship management.
REST Resource Layer
API endpoints inherit from EntityResource<E, R extends EntityRepository<E>>. This provides automatic CRUD operations, pagination, and field filtering without boilerplate code.
Frontend Type Safety
The UI enforces strict TypeScript discipline through generated types:
- Import all entity types from
openmetadata-ui/src/main/resources/ui/src/generated/ - Use components from
openmetadata-ui-core-componentsfor consistency - Localize strings through
useTranslation()with keys inlocale/languages/en-us.json
Python Ingestion Topology
Connectors implement a ServiceTopology composed of TopologyNode elements. This declarative pattern drives depth-first extraction and handles complex dependency graphs between metadata objects automatically.
Writing Tests Before Code
OpenMetadata enforces test-driven development (TDD) for every change. The CI pipeline blocks merges without adequate test coverage.
| Language | Test Framework | Command |
|---|---|---|
| Java | JUnit + integration tests | mvn verify |
| TypeScript | Jest (unit) + Playwright (E2E) | yarn test and yarn playwright:run |
| Python | pytest | make unit_ingestion or pytest |
For backend entities, create a *IT.java integration test that exercises the full REST flow:
public class MyEntityIT extends BaseEntityIT<MyEntity, CreateMyEntity> {
@Override
public CreateMyEntity createRequest(String name) {
return new CreateMyEntity().withName(name).withDescription("Demo");
}
@Override
public void validateCreatedEntity(MyEntity entity, CreateMyEntity request) {
assertEquals(request.getName(), entity.getName());
assertEquals(request.getDescription(), entity.getDescription());
}
}
Running Quality Checks and Submitting Your Contribution
Before opening a pull request, ensure all automated checks pass:
Java Quality Commands
mvn spotless:apply # Auto-format code
mvn verify # Run unit and integration tests
TypeScript Quality Commands
yarn lint:fix # Fix linting issues
yarn test # Jest unit tests
yarn playwright:run # End-to-end tests
Python Quality Commands
make py_format # Auto-format Python code
make lint # Run pylint and mypy
pytest # Unit tests
All checks are enforced in CI—pull requests cannot merge without passing the full pipeline.
Pull Request Workflow
- Fork the repository and create a feature branch:
git checkout -b my-feature - Commit using the message convention described in
CONTRIBUTING.md - Push your branch and open a GitHub pull request, linking any related issues
- Respond to reviewer feedback—the team expects tests, lint passes, and documentation updates for every change
Code Example: Adding a Backend Entity
Here's a complete example implementing a new entity in the OpenMetadata backend, following the patterns established in EntityResource and EntityRepository:
// src/main/java/org/openmetadata/service/resources/myentity/MyEntityResource.java
package org.openmetadata.service.resources.myentity;
import org.openmetadata.schema.entity.services.MyEntity;
import org.openmetadata.service.resources.Collection;
import org.openmetadata.service.resources.EntityResource;
import org.openmetadata.service.security.Authorizer;
import org.openmetadata.service.util.Limits;
import javax.ws.rs.*;
import javax.ws.rs.core.MediaType;
import org.eclipse.microprofile.openapi.annotations.tags.Tag;
@Path("/v1/myEntities")
@Tag(name = "MyEntities")
@Consumes(MediaType.APPLICATION_JSON)
@Produces(MediaType.APPLICATION_JSON)
@Collection(name = "myEntities")
public class MyEntityResource extends EntityResource<MyEntity, MyEntityRepository> {
public static final String COLLECTION_PATH = "/v1/myEntities/";
public static final String FIELDS = "owners,tags,domain";
public MyEntityResource(Authorizer authorizer, Limits limits) {
super(Entity.MY_ENTITY, authorizer, limits);
}
// CRUD endpoints inherited from EntityResource—no additional code required
public static class MyEntityList extends ResultList<MyEntity> {}
}
Key implementation points:
- Extends
EntityResourceto automatically expose CRUD, pagination, and field filtering - Declares
COLLECTION_PATHandFIELDSconstants as required by the base class - Uses
Entity.MY_ENTITYconstant registered inEntity.java - Annotates with
@Collectionfor automatic endpoint registration
Summary
- Set up the correct development environment for your target module: Java 21 for backend, Node 18+ for UI, Python 3.10-3.12 for ingestion
- Follow schema-first design by editing JSON schemas in
openmetadata-spec/and runningmake generate - Use established architectural patterns:
EntityResourceandEntityRepositoryfor backend, generated types for UI,ServiceTopologyfor Python connectors - Write tests before implementation using JUnit, Jest, or pytest depending on your module
- Pass all quality checks with
mvn verify,yarn test, ormake lintbefore submitting - Submit pull requests through the standard GitHub fork-and-branch workflow with proper commit conventions
Frequently Asked Questions
What programming languages does OpenMetadata use?
OpenMetadata is a polyglot codebase using Java 21 for the backend services, TypeScript/React for the UI, and Python 3.10-3.12 for the metadata ingestion framework. Contributors typically specialize in one module but should understand how JSON schemas tie all three together.
Do I need to write tests for my contribution?
Yes—test-driven development is mandatory. The CI pipeline enforces this through /test-enforcement policies. Java contributions need JUnit integration tests (*IT.java), TypeScript requires Jest unit tests and Playwright E2E tests, and Python uses pytest. Pull requests fail without adequate coverage.
How does the schema-first design work?
All entities are defined as JSON schemas in openmetadata-spec/src/main/resources/json/schema/. When you modify a schema, running make generate automatically produces Java POJOs, Python models, and TypeScript types. This ensures type consistency across all modules without manual synchronization.
What makes a good first contribution?
Start with issues labeled "good-first-issue" on the GitHub tracker. These range from documentation fixes to small bug fixes. Alternatively, adding a new connector follows a well-documented checklist in DEVELOPER.md (lines 71-82) and provides structured guidance for first-time contributors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →